Make it better with your history
No data: your model is ready in minutes. With your decisions: the same builder makes it better. On 3,080 real bank messages, models built with no data scored 84.6–86.8% against Jev's 86.6%; after about 400 reviewed cases, 91.7%; with the bank's history, about 95% (see Bank support, from day one). Past cases with the answers you gave them are used automatically, in any build of a model made with POST /v1/build: its first build, a rebuild under the same name, or a retrain.
Send your past cases
Two ways, same result:
- In the request: add
historytoPOST /v1/build, up to 50,000 cases. - As an upload: upload them to the model (
POST /v1/domains/{domain}/data, a CSV, JSONL or JSON file, see Upload history), then build again under the same name or retrain.
Each case is {"state", "answers"}, the state exactly as your program sends it, or a flat row of fields and answers, exactly as in an upload.
curl https://api.canonopylabs.com/v1/build \
-H "Authorization: Bearer $CANONOPY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": { "order": { "amount": 42.5 }, "customer": { "tier": "pro", "fraud_flag": false } },
"questions": { "action": { "type": "choice", "instructions": "What should we do with this refund request?",
"criteria": { "approve": "refund the customer", "escalate": "a person decides",
"decline": "politely decline" } } },
"name": "refunds",
"history": [
{ "state": { "order": { "amount": 820 }, "customer": { "tier": "pro", "fraud_flag": false } },
"answers": { "action": "escalate" } },
{ "state": { "order": { "amount": 35 }, "customer": { "tier": "free", "fraud_flag": false } },
"answers": { "action": "approve" } }
]
}'understood.history says how many were stored (accepted, rejected, and the problems with the first few). They're stored as one upload of the model: listed, replaceable and deletable like any upload.
What your history changes
- Your cases are the main examples. Your model learns from your real cases and their answers first. For text fields, it also learns from variations of your messages: typos, slang, other wording.
- Situations fill the gaps. Answers your history barely has, and rare cases it doesn't cover, are filled with situations for your questions.
history.gapslists them:{"question": "action", "answer": "decline", "your_cases": 12, "situations": 900}. - Rules found in your history are proposed, not used. Each comes with how many past cases it applies to and how often they agree with it. You confirm or reject each one; only confirmed rules are used. See Rules found in your history.
- Cases that contradict your rules are flagged. A past case that goes against one of your hard rules, or a rule found in your history that you confirmed, is never learned silently: it isn't learned at all until you review it. See Flagged past cases.
- It's measured on your own cases. About 15% of your past cases (at most 2,000) are held back, never learned from, and your model is measured on them. That's your honest real-world number.
Measured on your own held-back cases
After the usual quality check on fresh situations, your model is compared with the answers you gave your held-back cases:
"own_cases": { "cases": 600,
"questions": [{ "question": "action", "agreement": 0.962, "cases": 600 }],
"statement": "Measured on your own cases: your model agrees with your past answers in 96.2% or more of 600 held-back cases it never learned from, for every question." }- It's in the build's
quality.own_cases, and the version's report shows the same numbers withmeasured_on: "your_cases". - The same cases are held back on every retrain (plus a share of your new ones), so versions are compared on the same cases.
- With fewer than 40 past cases nothing is held back, and
own_casesisnull. - Serving still needs the usual 95% bar on fresh situations. Where your past answers differ from your instructions and rules, the number on your own cases shows it.
How your cases were used
The build's history:
"history": { "cases": 4200, "learned_from": 3558, "held_back": 630, "variations": 2400,
"flagged": 12, "pending": 12,
"gaps": [{ "question": "action", "answer": "decline", "your_cases": 12, "situations": 900 }],
"statement": "…" }| field | meaning |
|---|---|
cases | your past cases with an answer |
learned_from | the ones your model learns from: not held back, and not flagged and waiting for you |
held_back | held back to measure your model on your own cases |
variations | variations of your messages it also learns from |
flagged, pending | cases that contradict a rule, and those still waiting for your review |
gaps | answers your history rarely has, filled with situations |
The flow
- Build with
history(or upload, then build again or retrain). Your model trains, is checked on fresh situations and measured on your own cases, with nothing to do in between. - Optional, once it's ready: confirm or reject the rules found, review the flagged cases (until then they aren't learned), and review how it decides on about 20 examples. Your decisions and corrections apply at the next retrain.
- Later, upload more cases and retrain: the same held-back cases, so you see real progress.
In the console, the build page shows the rules found and the flagged cases, and the model page's Review examples shows how it decides. With a coding agent, the MCP tools are build_model (with history rows, or history_file on the local server), get_build, confirm_found_rules and review_flagged_history. In the CLI: canonopy build … --history FILE, canonopy rules found and canonopy history flagged.