Skip to content
Docs / Outcomes and options

Improving

Each of these is something you do once; the next version takes about a minute. Every training report ranks them for you.

For a model made from a Jev request (POST /v1/build), also see Improve your model: advice on what would help most, targeted retrains with a before and after, rule changes that retrain straight away and say what they change, and marking a decision wrong (which works for every model).

See what's waiting, then retrain

GET/v1/domains/{domain}/progress

Your model learns from new cases only when you retrain it: nothing retrains on its own, and nothing is sent to you. The console's Progress tab (and canonopy progress refunds) shows what the next retrain would learn from, since the latest version:

  • uploaded past cases, reported outcomes, unsure cases you (or our backup) answered, cases synced from offline devices, examples you corrected, and decisions you marked wrong;
  • how many of the new outcomes differ from the answer it gave: those teach it most.

It says it in one line: "48 new cases since version 4, 12 where the real outcome differed from its answer. Retrain to learn from them." Then retrain with the button, canonopy train refunds or POST /v1/domains/{domain}/train. The counts start again from the new version.

It also shows advice: what would improve your model most right now.

The same view lists every version: when it was made, how many of your cases it learned from, its accuracy per question and the change against the version before. When the report compared both versions on the same held-back cases, that is the number shown; otherwise each version was measured on its own held-back cases, and it says so. A starter model is measured on starter example cases, not your own, and is marked as such. Once a version has 20 reported outcomes, you also see how often its answers matched them.

Report outcomes

POST/v1/outcomes

Tell us what really happened on a decision. Outcomes feed the next version and day-one graduation.

json
{ "decision_id": "dec_4f1c2a9b0e7d6a51", "action": "block-card", "answers": { "urgency": 2, "needs_human": false } }
json
{ "recorded": true, "decision_id": "dec_4f1c2a9b0e7d6a51",
  "graduation": { "urgency": { "status": "fallback", "version": null, "outcomes": 213, "needed": 787 } } }

action is shorthand for the decision question's answer.

The unsure queue

GET/v1/domains/{domain}/unsure?limit=50
POST/v1/domains/{domain}/answers

Decisions held for a person wait here until someone answers them, in the console or over the API: review (still below the model's safety bar after the message was double-checked) and ask (your rules allow no action).

json
{ "decision_id": "dec_7a…", "answers": { "action": "route-payments" } }

Every answer is captured, and the next version learns from it, so it keeps getting better on your own traffic. Your own LLM or Jev can answer the queue too: see Using with LLMs.

Add an option or an action

POST/v1/domains/{domain}/options

A new Choice option or action in about 30 seconds, without breaking the others:

json
{ "question": "action", "option": "freeze-account",
  "description": "freeze the whole account, including the app and all cards",
  "when": "topic is lost_stolen or unrecognised and new_device_login_24h is true",
  "never_when": [{ "field": "new_device_login_24h", "op": "is_false" }] }

The job's report has a before/after:

json
{ "question": "action", "option": "freeze-account", "took": "under a minute",
  "before": { "accuracy": 0.958, "cases": 988 },
  "after":  { "accuracy": 0.955, "cases": 988, "option_recall": 0.93, "option_cases": 58, "unchanged_cases_accuracy": 0.957 } }

In the bank test, adding freeze-account took 29 seconds and left the other nine actions at 96.0%.

Clarify a rule where it hesitates

Most mistakes sit on blurry boundaries, like "is this a card question or a payments question?". The report shows those cases; say once, in plain words, which way each goes (POST /v1/domains/{domain}/agent), and train again. On a model made with POST /v1/build, a rule change shows what it changes first: see Change a rule with a preview.

Upload more of what's weak

If lost or stolen cards are 89% right and climbing with data, the report estimates what 500 more such cases would add. Upload them and train.

Replace cases that taught a bad habit

If some of your past cases were resolved badly (old recordings, a policy you've since changed), don't start a new model: replace them and train the same one. Your code keeps calling refunds@latest, and the next version is refunds@2.

bash
curl "https://api.canonopylabs.com/v1/domains/refunds/data?mode=replace" \
  -H "Authorization: Bearer $CANONOPY_API_KEY" -F file=@better-history.csv
curl -X POST https://api.canonopylabs.com/v1/domains/refunds/train \
  -H "Authorization: Bearer $CANONOPY_API_KEY" -H "Content-Type: application/json" -d '{"wait": true}'

The new version is compared with the serving one on your current held-back cases, with their current answers, so better cases can win. To remove just one upload instead, list them (GET /v1/domains/{domain}/uploads) and delete it. Deleting is permanent; versions already trained are kept. See Replace or remove cases.