Skip to content
Docs / Training and reports

Training and reports

Everything that makes a model runs on our servers. You send cases and get back a version with a report. A new version usually takes about a minute.

The quickest start needs no past cases: build a model from the JSON you send Jev, see Start with no data. This page covers the other way in, describing the decision to the set-up agent, and what every version's report tells you. A model made that way uses your past cases as described in Make it better with your history, and its retrains are targeted retrains.

Set up the domain

POST/v1/domains

Most people start in the console: New decision model opens a conversation with the set-up agent. The same thing over the API:

POST/v1/domains/{domain}/agent
json
{ "message": "We handle refund requests. Actions: approve, escalate, decline. Fields: message (text), amount (number), tier (category: free, pro, enterprise), fraud_flag (yes/no). Never approve when amount is over 500. Always escalate when fraud_flag is true." }

The reply tells you what it set up, in plain words, and what to do next (next_step: describe while it still has questions for you, then train). For example: "Set up 3 actions, 4 state fields and 2 rules that can't be broken. Next, train it (looking at the example cases is optional, any time)." You can train right away. Uploading past cases first is optional, and so is reviewing the examples.

Upload history

POST/v1/domains/{domain}/data

Each past case and how it was resolved: the outcome you already record is enough. Send a multipart/form-data file (.csv or .jsonl), a text/csv or application/x-ndjson body, or JSON {"rows": [...]}. A row is either {"state": {...}, "answers": {"question": value}} or a flat row whose columns are state fields and question ids.

bash
curl https://api.canonopylabs.com/v1/domains/refunds/data \
  -H "Authorization: Bearer $CANONOPY_API_KEY" \
  -F file=@history.csv
json
{ "accepted": 9800, "rejected": 12, "total_cases": 9800, "per_question": { "action": 9800 },
  "problems": [{ "row": 17, "problem": "action: 'refnd' is not a valid answer (options: approve, escalate, decline)" }] }

Every row is checked. A row with a number that isn't a number, an answer that isn't one of the options, or a question the model doesn't have is skipped, counted in rejected and listed with its row number (the first 20). Empty values, NaN and infinity count as missing. Uploads can be up to 50 MB; split bigger files.

Your answers are the truth. Where your history says how a case was resolved, that is what your model learns and is measured against. Your rules still can't be broken: a past answer that one of them doesn't allow is counted in the report (fix_history).

On a model made with POST /v1/build, uploaded cases are used as in Make it better with your history: rules found in them are proposed for you to confirm, and cases that contradict your rules wait for your review.

Each upload is kept as one unit with an id (upload_id in the answer), so you can replace or remove it later. Add ?file_name=history.csv to name a raw CSV or JSONL body in the upload list; a multipart upload uses the file's own name.

Replace or remove cases

Improve the model you already have instead of starting a new one. Your code keeps calling refunds@latest; the next training run makes refunds@2, refunds@3, and so on, from the cases the model has at that moment.

GET/v1/domains/{domain}/uploads

Every upload, newest first: its id, date, file name, how many cases it holds, and how many rows were set aside. Examples you corrected are listed too (source: "signoff"). Cases uploaded before uploads had ids appear as one group with the id earlier.

json
{ "domain": "refunds", "total_cases": 6361, "from_uploads": 6349, "from_signoff": 2, "from_outcomes": 12,
  "uploads": [
    { "id": "upl_4f1c2a9b0d3e5f67", "source": "upload", "mode": "replace", "file_name": "cases-sep.csv",
      "created_at": "2026-09-27T09:12:00Z", "cases": 6347, "rejected": 3 },
    { "id": "upl_0a1b2c3d4e5f6a7b", "source": "signoff", "mode": "add", "file_name": null,
      "created_at": "2026-09-20T15:02:00Z", "cases": 2, "rejected": 0 } ] }

Replace every uploaded case with a new file, in one step:

bash
curl "https://api.canonopylabs.com/v1/domains/refunds/data?mode=replace" \
  -H "Authorization: Bearer $CANONOPY_API_KEY" \
  -F file=@better-history.csv

The earlier uploaded cases are deleted once the new file's rows are in, and replaced says how many went. If no row of the new file is accepted, nothing is stored and nothing is removed.

Delete one upload:

DELETE/v1/domains/{domain}/uploads/{upload_id}

Delete every uploaded case. So it can't happen by accident, confirm with the model's name:

DELETE/v1/domains/{domain}/data?confirm={domain}
json
{ "domain": "refunds", "removed": { "uploaded_cases": 6278, "signoff_cases": 0, "decisions": 0 },
  "uploads_removed": ["upl_9e8d7c6b5a4f3e2d"],
  "remaining": { "total_cases": 14, "from_uploads": 0, "from_signoff": 2, "from_outcomes": 12 },
  "message": "Deleted permanently: 6,278 uploaded cases. Train again to make the next version from the cases left. Versions already trained are kept: they're yours, and still download and serve." }

What each call removes:

callremoveskeeps
POST …/data?mode=replaceevery uploaded case, once the new file's rows are inexamples you corrected, outcomes
DELETE …/uploads/{upload_id}that upload's cases (or a review's corrected examples, when you name them)everything else
DELETE …/data?confirm=…every uploaded caseexamples you corrected, outcomes
…&signoff=truealso the examples you correctedthe review record: who reviewed, when, and which examples
…&outcomes=truealso the hosted decisions the next version would learn from: those with a reported outcome or an unsure-queue answer, and those our hosted backup answeredother hosted decisions
…&starter=true, or DELETE …/uploads/starteralso the starter cases of a question that reads text (they then stay off for this model)everything else
  • Deleting is permanent. The cases are removed from our database, not hidden. Backups roll off on their usual schedule.
  • Versions already trained stay yours. They keep serving and downloading; roll back to one any time.
  • Examples you corrected stay unless you remove them. They are answers you gave yourself, so replacing or clearing uploads leaves them in place.
  • While a new version is being made, replacing and deleting answer 409: wait for the job (GET /v1/jobs/{id}), then try again. Adding cases is always fine.

Then train again. The report's cases says what changed ("Trained on 6,349 cases; 6,278 earlier cases were removed since version 3"), and the comparison with the serving version uses the current held-back cases with their current answers. So new cases that fix a bad habit can pass it, and a version that is really worse on your current cases is still kept out.

The console does the same on the model's Data tab: the upload list with a delete button per upload, Replace all cases when you upload a file, and Clear all after you type the model's name.

Train

POST/v1/domains/{domain}/train
json
{ "promote": "auto", "wait": true }
  • promote: "auto" promotes the new version when its report recommends it; always and never override that.
  • With wait: false you get a job to poll at GET /v1/jobs/{id}.
  • It learns from the cases the model has now. After a replace or a delete, that means the current cases only.

The answer is a job with the report attached:

json
{ "id": "job_0a1b2c3d4e5f6a7b", "domain": "bank-support", "kind": "train", "status": "done",
  "version": 3, "message": "Version 3 is serving.", "report": { "…": "…" } }

Review the examples (optional)

GET/v1/domains/{domain}/examples
POST/v1/domains/{domain}/signoff

About 20 example cases, each with the answer your setup gives and why. Look at them any time, before or after training: nothing waits on them. If you have history, upload it first: the examples are then drawn from your own cases.

To correct one, send only the ones you correct. "retrain": true starts the retrain in the same call, and the answer has its job_id and a message:

json
{ "examples": [
    { "id": "ex_02", "ok": false, "correct": { "action": "escalate" }, "note": "pro accounts under 30 days go to a person" } ],
  "retrain": true }

Without retrain, your corrections are learned at the next retrain. The review is kept as a record. In the console, it's Review examples on the model's page; in the CLI, canonopy examples refunds and canonopy signoff refunds --correct ex_02:action=escalate --retrain.

The report

GET/v1/domains/{domain}/versions/{version}/report

Every number comes from your own held-back cases, which are never used for learning. The one exception is a starter model for a question that reads text, before you have enough cases of your own: its accuracy is measured on starter example cases, marked measured_on: "starter_cases" and explained in the report's starter section. A model you built with no data is the same: until you have enough answered cases of your own (about 200), it is measured on held-back examples of situations, marked measured_on: "situations". From then on those examples only teach (at full weight, they never fade), and the accuracy, the calibration and the never-worse check use your own held-back cases only; the gate's measured_on says which. The console shows it as cards; the JSON has four parts.

1. Summary

  • accuracy per question, against the previous version and your day-one backend (previous, day_one, change_points);
  • how it did on rare situations it was tested on;
  • calibration per question, with a plain sentence ("When it says 90% sure, it's right 89% of the time");
  • how many cases each rule matched and how many decisions it changed. Violations are always 0;
  • a promote or keep recommendation, with the reason.

cases says how many of your cases the version learned from and, when cases were deleted since the last version, how many. starter (when a question reads text and has few of your own cases) says whether a starter model is answering, still helping some options, or has stepped aside.

2. Weak spots

  • the weakest options, with their accuracy and case counts;
  • the most-confused pairs, with real example cases;
  • for day-one questions: how many more outcomes until they graduate.

3. What would make it better

A ranked list, each item with an estimated gain and a fix: the exact API call that acts on it, so the console can offer it as one click.

kindWhat it suggests
more_casesmore cases where accuracy is still climbing, estimated from your own learning curve
clarify_rulea boundary your rules don't separate: say once which way the example cases go
add_optionan option for cases that fit none
more_language_casesa language below the bar: "upload ~300 cases in Swahili to handle more automatically"
fix_historypast cases whose answer one of your rules doesn't allow: your rule was followed for them; change the rule if it's too strict

4. The safety bar

The bar this version set for itself, in plain words: "automatic answers 97%+ accurate; 78% handled automatically; the rest double-checked". For each question: the confidence an answer needs, the share of your held-back cases handled automatically, and how often those were right. Then per language: the strong languages (answered automatically) and the ones still double-checked. There's no threshold to set; each retrain measures it again. See Languages.

Versions

GET/v1/domains/{domain}/versions
POST/v1/domains/{domain}/promote

@latest serves whichever version you promoted. A roll-back is a promote of an older version: {"version": 2}.

To see how accuracy moved from version to version, and what the next retrain would learn from, open the model's Progress tab or call GET /v1/domains/{domain}/progress. See See what's waiting, then retrain.