---
name: canonopy-decisions
description: Use Canonopy Decisions (canonopylabs.com) to give a decision that is made over and over, with a fixed set of possible answers, its own decision model. Covers ticket triage and routing, refunds and approvals, moderation, fraud flags, scoring, guardrails, model routing, and game or agent loops. Use it to build a model from the JSON sent to Jev (no data needed), improve it with past cases and rules found in them, set up, train, retrain and call a Canonopy decision model through its MCP tools or HTTP API, to move code from TypeSafe's Jev (same request format), to run a downloaded model offline, and to answer questions about Canonopy such as pricing, the trial, answer sources and routes, errors and limits.
---

# Canonopy Decisions

Canonopy Decisions gives a team **its own model for a decision it makes over and over**: route a ticket, approve a
refund, flag a transaction, pick an agent's next action. The pitch: *"Send us the JSON you send Jev. In minutes, get
your own 1 MB model. For games and rule-based decisions, it beats Jev from day one. For text, it beats Jev clearly
once it learns from your decisions."* No data is needed to start: it's built from the request the team already sends
Jev, and their decisions (answers to unsure cases, past cases) make it better. Their rules are enforced on every
decision, and it answers in the same request and response format as TypeSafe's Jev (System One: Choice, Score and
Noul questions). It runs on Canonopy's servers or, once downloaded, on the team's own machines, even offline.

A decision model is called a **domain** in the API (`/v1/domains`). Its name is what `decide` takes as `model`:
`refunds@latest`, or a fixed version like `refunds@3`.

Supporting files, read only when needed:
- `reference/api.md`: every endpoint with its MCP tool, request and answer shapes, the report's fields, errors and limits.
- `reference/build.md`: building from a Jev request, past cases, rules found in them, flagged cases, advice, targeted
  retrains, rule changes with a preview, and marking decisions wrong: every field and call.
- `reference/describing.md`: how to describe a decision so the set-up turns every sentence into a rule.
- `reference/troubleshooting.md`: 402, 409, 413, low accuracy, too many reviews, refused addresses, the offline queue.
- `reference/offline.md`: download a version, run it offline, your own fallback, save and sync later.
- `examples/refunds-walkthrough.md`: one decision model from nothing to production, tool call by tool call.
- `examples/switch-from-jev.md`: moving Jev code over, in curl, Python and JavaScript.

## When Canonopy fits

Good fits are **high-volume, bounded decisions**: the same question asked many times, answered from a fixed set of
options or levels.
- triage and routing (support tickets, emails, claims, alerts);
- approvals and actions (refunds, credit, escalations), moderation, fraud and risk flags;
- scoring (priority, urgency, risk) and guardrails (check an action an LLM proposes before it runs);
- model routing (which model or handler takes a request), and game or agent loops (the next move or tool).

It is **not** the tool for:
- one-off questions nobody will ask twice, or open-ended questions without a fixed answer set;
- writing prose, replies or summaries. Pair it with an LLM instead: call `decide` first, then have the LLM write the
  reply for `decision.action` only (so it can't promise what the rules block), and hand `ask` cases to a person.
  Patterns: https://canonopylabs.com/docs/using-with-llms.md.

When a question is new or still changing, offer the **day-one backend**: the team's own Jev key, an
OpenAI-compatible model, or Canonopy's built-in model answers it in the same shape (`source: "fallback"`), with their
rules still applied, until the question has enough real outcomes to move to their own model (step 12 below).

What to bring: **nothing but their Jev request** (one real example of the state, the questions as they are, and
rules in plain words if they like). That's the same for text, data and games. **Their decisions make it better:** past
cases and how each was resolved (the outcome or category already recorded is enough), answers to unsure cases, or
recorded play for a game. Offer it once the model is working, not as a condition to start.

## Connect

Pick one. The key always comes from the console (https://canonopylabs.com/console/keys) and lives in configuration or
an environment variable, never in the chat.

**Hosted MCP connector** (nothing to install; Streamable HTTP):
```bash
claude mcp add --transport http canonopy https://api.canonopylabs.com/mcp --header "Authorization: Bearer $CANONOPY_API_KEY"
```
Add `--scope user` for every project. Cursor, Codex and other clients: https://canonopylabs.com/docs/coding-agents.md.
Hosted, the tools can't read or write files on the user's machine: `upload_history` and `replace_history` take the
cases inline (`rows`, or a file's contents as `text`, up to about 2 MB per call), `download_model` returns a link that
works for 10 minutes, and a day-one backend's key is added in the console (open the decision model, then **Day one**).

**Local MCP server** (uploads files from disk, saves downloads to disk, reads a backend key from an environment
variable; Python 3.9+):
```bash
pip install https://canonopylabs.com/dl/canonopy_cli-0.6.2-py3-none-any.whl
claude mcp add canonopy --env CANONOPY_API_KEY="$CANONOPY_API_KEY" -- canonopy mcp
```

**Plain HTTP** (no package needed): base URL `https://api.canonopylabs.com`, header
`Authorization: Bearer $CANONOPY_API_KEY`. Decisions go to `POST /v1/decide`, or its Jev-compatible alias
`POST /v1/systemone` (same body). The API describes itself at https://api.canonopylabs.com/openapi.json.

Other packages, all optional, served from https://canonopylabs.com/dl/:
- the runtime, to run a downloaded model: `pip install "canonopy-runtime[text] @ https://canonopylabs.com/dl/canonopy_runtime-0.6.3-py3-none-any.whl"`
- the Python SDK (`from canonopy import Client`; Jev-style `Choice`, `Score`, `Noul`): `pip install https://canonopylabs.com/dl/canonopy_decisions-0.4.0-py3-none-any.whl`

Each MCP tool call is a normal API call with the same key, limits and plan; nothing is billed per call.

## Start from a Jev request (no data needed)

The fastest way in, and the one to offer first when the user already calls Jev or has a Jev request (`{"model",
"state", "questions"}`) in their code: **send the same JSON to `build_model`** (HTTP `POST /v1/build`, alias
`POST /v1/systemone/build`) and get their own model back in minutes. No past decisions or answers needed; a few
examples of what you send Jev make it better. Full detail: `reference/build.md`.

0. **JSON first.** Nothing is built without the user's own request (one real `state` and their `questions`).
   Ask them: "Paste the request you send Jev." If they don't have one, find where the decision is made in their
   code (or where Jev is called), take `state` from the real object the program has at that moment and `questions`
   from the decision's possible outcomes, then show the JSON and get their yes before sending it.
   `draft_request_help` gives the shape and a checklist; `build_model` without the JSON says the same and builds
   nothing.
1. **Build.** `build_model` with `state` (ONE real example, text or the JSON object their program sends),
   `questions` (exactly as sent to Jev), and optionally `examples` (recommended: 5-20 more real states like `state`,
   up to 50, states only; ask the user for a few, or take them from their logs or tests), `fields` (what fields
   mean, e.g. `{"order.amount": "the refund asked for, in USD (0-5,000)"}`), `rules` (plain words), `name`,
   `description` and `history` (past cases as rows; the local server also takes `history_file`). A pasted Jev
   request works as is (`model` is ignored). Read `understood` back to the user: the fields read (`fields`), what
   isn't read (`not_read`), the notes used (`field_notes`) and not matched (`field_notes_not_matched`), the rules
   enforced (`rules`) and any `rules_not_understood`, with the reason.
   Through MCP it asks a few questions about the JSON first (`interview`, default true; the build waits up to 10
   minutes, then goes on with what it has). `build_model` returns when the model is ready, or sooner when its
   questions wait for answers. Nothing else waits for the user.
2. **Answer the questions.** `get_build_questions` (HTTP `GET /v1/build/{build_id}/questions`): at most about 25,
   most useful first, each about a field, question or choice of the JSON (`refs`) with the keys its answer takes
   (`answer_with`): what fields mean, units and ranges, whether a yes/no field is computed by their program, how to
   tell two choices apart, 10-50 real examples across named situations, unwritten priorities. Answer from their
   code, config and logs; ask the user when unsure; skip what you can't answer. Only real facts and real states,
   never guesses. `answer_build_questions` (`answers`: each `id` plus `text`, `unit` and `range`, `ready_made`,
   `examples`, or `skip`; `done: true`); the build goes on at once. `rejected` says why an answer wasn't recorded.
   A short follow-up (round 2) comes only when the check finds a gap; its answers are used at the next retrain.
3. **Follow it.** `get_build` (`build_id`, or `model` for the latest build; HTTP `GET /v1/build/{build_id}` or
   `GET /v1/domains/{domain}/build`) until `status` is `ready`, `needs_attention` or `failed`. Through MCP, pass
   `wait: true`; if it returns `still_working`, call it again with `wait` (nothing is needed from the user). Statuses:
   `queued`, `building`, `training`, `checking`; phases: `understanding`, `questions`, `preparing`, `training`,
   `checking`. Quote `message` and `next_step`. Tell the user: "Your model is ready in minutes. Review how it decides
   any time (optional)."
4. **The quality check.** Each question must agree with the user's instructions and rules in at least 95% of fresh
   situations (`quality.bar`). `ready`: it serves as `<name>@latest`. `needs_attention`: not serving; quote each
   question's `advice`, then clarify the instructions or rules and build again under the same name, or retrain.
5. **Try it on real work, then switch over.** Games: the user runs `canonopy playtest <model>@latest --cmd "<their
   game>" --episodes 20` on their machine (their program prints one JSON line of numbers per episode; only those
   numbers come back; `--keep-situations N` lets the next retrain learn where their game really goes). Everything
   else: `shadow_mode` (start), then `send_shadow_cases` with real requests and the decision they made today; the
   model decides silently, and `review_shadow_disagreements` lists where it disagreed (at most 20): show them to
   the user and record only their review. Then switch over: in their code, change the base URL to `https://api.canonopylabs.com` and `model` to
   `<name>@latest`; they keep sending the same JSON to `/v1/systemone` or `/v1/decide`.

**How it decides (optional, any time once it's ready):** if the user wants to see it, `get_examples` (also the
build's `examples`) and, for any they say is wrong, `correct_examples` (see "6. Review how it decides" below).

In Python, `@canonopy.fn` over a typed function does the same (swap `@jev.fn` for `@canonopy.fn`); the CLI is
`canonopy build jev-request.json`. Limits per workspace: 5 builds a day, 3 during the free trial, 30 a month with a
subscription (`429`, or `402` when the trial's builds are used up); a build is a new model, a new version or a rule
change, and retrains don't count (fair use: 20 a day). A build counts as training (`402` after the trial without a
subscription). `usage` shows where the workspace stands.

**With past cases** (in `history`, or uploaded with `upload_history` and then built again or retrained), the same
build uses them automatically: their cases are the main examples, and:
- **Rules found in their history** (`found_rules`): each with `support`, `agreement` and `changes` (what confirming it
  changes). Only confirmed ones are used. Show each `statement` and `changes.statement` and let the user decide;
  never confirm on their behalf: past habits get copied too. `confirm_found_rules` (`rules`: each `id`,
  `decision`: `confirm` or `reject`, and `hard: true` only when `can_be_hard` and the user wants it enforced on every
  decision). HTTP `POST /v1/build/{build_id}/found-rules`.
- **Flagged past cases** (`flagged`): cases that contradict a hard rule or a confirmed found rule. Not learned at all
  until reviewed. `review_flagged_history` lists the pending ones; decide with the user: `follow_rule`, `keep` or
  `drop` per case (`rows`), or `all`. HTTP `GET` / `POST /v1/build/{build_id}/flagged`.
- **Measured on their own cases:** `quality.own_cases.statement`, on about 15% of their cases held back (the same
  ones on every retrain). Quote it: it's the real-world number.

**Spot checks** (once it acts): a few confident decisions are set aside at random for a person, on one optional
weekly card (`spot_checks` with `model`; HTTP `GET /v1/domains/{domain}/spot-checks`): `light` (about 1 in 100, the
default), `thorough` (2 in 100) or `off` (ask the user before changing it). Answer with `answer_unsure`. Progress
shows "spot-checked accuracy". What the model learns counts real outcomes most, then a person's answers (reviews,
spot checks, shadow reviews), then confident Jev answers least.

**Improving a built model** (every retrain is the user's call; nothing is automatic):
- `get_advice` (`GET /v1/domains/{domain}/advice`, also in `get_progress`): `retrain_to_sharpen`,
  `rule_may_be_wrong`, `answer_these`, `upload_these`, `answer_questions` (the follow-up), `shadow_review`,
  `from_playtest`. Quote `text`; offer the `action`.
- **Targeted retrain:** `train` on a built model returns a `build_id`; `get_build` shows `retrain.statement` with the
  before and after per area.
- **A rule is wrong:** `change_rules` (`model`, `rules`, `mode`: `add` or `replace`; HTTP
  `POST /v1/domains/{domain}/rules`). It retrains straight away; quote `rule_preview.statement` ("Changes 1,240 of
  3,000 situations: approve → escalate …"). Its about 10 examples (situations whose answer changes) are there if the
  user wants to see them. The new version serves only if it passes the quality check.
- **A decision was wrong:** `mark_decision_wrong` (`decision_id`, `answers` or `action`, `note`; HTTP
  `POST /v1/decisions/{decision_id}/wrong`). Recorded as its outcome; the next retrain learns it.

## The workflow

The set-up way, for any decision model (and the steps a built model shares). Each step gives the MCP tool first,
then the HTTP call. `{domain}` is the decision model's name.

### 1. Create a decision model

`create_decision_model` with a `name` (lowercase letters, digits and dashes) and a one-line `description`; pass
`describe` to do step 2 in the same call. HTTP: `POST /v1/domains`. A taken name is a `409`: check first with
`list_decision_models`.

### 2. Describe it in plain words

`describe_decision_model` with a `message`. HTTP: `POST /v1/domains/{domain}/agent` with `{"message": "..."}`.
The reply has `reply`, `summary` (the setup so far, in plain words), `open_questions` and `next_step`
(`describe`, `train`, or `follow_build` when a message to a built model started a rule change; older set-ups may say `review_examples` or `upload_history`, both optional).
Call it again while it's `describe`; once it's `train`, train right away (step 4). Uploading past cases first (step 3)
is optional.

Write it the way the set-up reads best (full guide: `reference/describing.md`):
- **Actions:** `Actions: approve, escalate, decline.`
- **Fields**, each with its type: `Fields: message (text), amount (number), tier (category: free, pro, enterprise), fraud_flag (yes/no).`
- **Never** and **always** rules, one per sentence, on fields: `Never approve when amount is over 500; escalate instead.`
  `Always escalate when fraud_flag is true.`
- **Only:** `When tier is free, only decline or escalate.` (the only choices then) or
  `Approve only when identity_verified is true.` (blocked otherwise).
- **Unless:** `Never approve when amount is over 500 unless tier is enterprise.`
- **Otherwise:** `Otherwise approve.` (what happens when no rule applies; without it the set-up asks).
- Comparisons it reads: is, is not, is over / above / more than, is under / below, at least, at most, is one of,
  contains; values like 500, $500, 2k, true/false, yes/no.

When the reply says **"I couldn't turn ... into a rule"**, the sentence was not used. Read the reason it gives and
say it again the way it suggests: add a missing field first (`Fields: angry (yes/no)`), use one of the listed
actions, or split a long sentence into one rule per sentence. A rule about what a message is *about* works when the domain has a question for it: "Always block-card when
topic is lost_stolen" is enforced whenever there's a real chance it applies, and unsure cases go to review
(`decision.rule_chances` shows the chance). Tell the user which sentences weren't used; never drop them silently.

### 3. Upload history (optional)

`upload_history`. HTTP: `POST /v1/domains/{domain}/data`.
- **Formats:** CSV with a header row, JSONL (one case per line), or JSON `{"rows": [...]}`. A row is
  `{"state": {...}, "answers": {"action": "approve"}}`, or flat: state fields as columns plus one column per question,
  named like the question, holding how the case was resolved. Values: an option name (Choice), a level number (Score),
  true/false (Noul).
- **Answers are optional for decisions on data:** upload the cases with their fields; where a row does record how
  the case was resolved, that answer is what the model learns. For decisions on text, the recorded answers are how it
  learns the team's categories, so include them. A row that can't be used is rejected with its reason (for example
  "no resolved answer for any question": add the answer column for that row).
- Every row is checked: the answer says `accepted`, `rejected` and the first 20 `problems` with row numbers. Fix and
  re-upload the rejected rows. Uploads up to 50 MB per request over HTTP (20 MB in the console, about 2 MB per hosted
  MCP call).
- **Each upload is a unit** with an `upload_id`. `list_uploads` shows them (`GET /v1/domains/{domain}/uploads`).
  `replace_history` swaps every uploaded case for a new file (`POST /v1/domains/{domain}/data?mode=replace`);
  `delete_upload` removes one (`DELETE /v1/domains/{domain}/uploads/{upload_id}`); `clear_history` removes every
  uploaded case (`DELETE /v1/domains/{domain}/data?confirm={domain}`). All three **delete permanently**: ask the user
  first. The MCP tools refuse to run until `confirm` is set to the model's name. Over HTTP only clearing takes
  `?confirm=`; deleting one upload and `mode=replace` act at once, so the agent's own confirmation is the only guard. Corrected examples and reported outcomes stay unless asked for;
  versions already trained are kept. `history_summary` counts the cases.

### 4. Train

`train`. HTTP: `POST /v1/domains/{domain}/train` with `{"promote": "auto", "wait": true}`. About a minute on
Canonopy's servers; the answer is a job with the report. `promote`: `auto` (serve it when the report recommends it),
`always` or `never`. With `wait: false`, follow the job with `get_job` (`GET /v1/jobs/{job_id}`). There are no other
settings to tune.

### 5. Read the report

`get_report` (`name`, optional `version`; default: the serving version). HTTP:
`GET /v1/domains/{domain}/versions/{version}/report`. The numbers come from the team's **own held-back cases**. Tell
the user, briefly, and quote the report's own sentences rather than explaining them:
- **Accuracy** per question, against the previous version and the day-one backend.
- **Calibration:** the report's plain sentence ("When it says 90% sure, it's right 89% of the time").
- **Share handled automatically:** `safety_bar.statement`, e.g. "Automatic answers 97%+ accurate; 78% handled
  automatically; the rest double-checked." It's a 97% safety bar, set automatically; there is nothing to set.
- **Rules:** how many decisions each rule changed; violations are always 0.
- **Recommendation** (`promote` or `keep`, with its reason), and whether it passed the never-worse check.
- **The to-do list** (`improve`), ranked, each with the exact call in `fix`. Offer the top one or two.
- **If a question says `measured_on: "starter_cases"`**, its accuracy isn't measured on the user's own cases yet: say
  so, and suggest uploading past cases or reporting outcomes to get a real number.

### 6. Review how it decides (optional)

Optional, any time (before or after training); nothing waits on it. Offer it if the user wants to see how it decides.
`get_examples`. HTTP: `GET /v1/domains/{domain}/examples` (on a built model, also the build's `examples` once it's
`ready` or `needs_attention`). About 20 cases, each with its answer and why. For a decision made on what a message
says, `status` can be `preparing` for about a minute after set-up while its example messages are prepared: ask again
shortly. An example with `described: true` names the kind of message instead of quoting one.

If the user says one is wrong, `correct_examples` with only the corrected ones (`corrections`: `{id, correct:
{question: answer}, note}`; `retrain`, default true, retrains in the same call: follow its `build_id` with `get_build`, or its `job_id` with `get_job`).
HTTP: `POST /v1/domains/{domain}/signoff` with `{"examples": [{"id": "ex_03", "ok": false, "correct": {"action":
"escalate"}}], "retrain": true}` (answer: `build_id` for a built model, `job_id` otherwise, and `message`). Without
`retrain`, they apply at the next retrain. On a built model the corrections also make its instructions give those
answers. Send only answers the user gave; the review is kept as a record.

### 7. Promote or roll back

`list_versions` (`name`), then `promote_version` (`name`, `version`). HTTP: `GET /v1/domains/{domain}/versions`,
`POST /v1/domains/{domain}/promote` with `{"version": 2}`. `@latest` serves the promoted version. A version that failed
the never-worse check stays a `candidate`; promote it by hand only if the user chooses to. A roll-back is a promote of
an older version. Ask before promoting: it changes what production answers.

### 8. Decide

`decide` with `model` (`name@latest` or `name@3`) and `state`. HTTP: `POST /v1/decide`:
```json
{ "model": "refunds@latest",
  "state": { "message": "Refund my order please", "amount": 40, "tier": "pro", "fraud_flag": false } }
```
Leave out `questions` to ask every question the model knows, or send questions in Jev's format (Choice: 1 to 255
options; Score: 2 to 10 levels; Noul: a yes/no probability). A question the model doesn't have is answered by the
day-one backend and listed in `candidate_questions`; it joins the model only when added with `update_decision_model`
(`PATCH /v1/domains/{domain}`; `questions` there replaces them all, so send the full set). Keep the answer's `id`.

### 9. Read `source`, `route` and `decision`

Each answer has **`source`**, who answered:

| `source` | meaning |
|---|---|
| `trained` | their own model |
| `translated` | their own model, after the message was translated to English |
| `fallback` | their day-one backend (question not learned yet, or still unsure after translation) |
| `backup` | still below the bar and no day-one backend answered: Canonopy's hosted backup answered from the allowed options; their rules still applied; included in the price |
| `rule` | their rules left only one option |
| `local_fallback` | a downloaded model offline: their own fallback function answered |

Each answer also has **`route`**: `act` (automatic, at or above the model's safety bar) or `review` (held for a
person). A day-one answer below 0.8 confidence is `review` too; for a Noul that means an answer between 0.2 and 0.8.

**`decision`** (when the decision question was asked): `action` (the final action with the rules applied),
`confidence`, `blocked_by_rules` (rule ids that blocked something), `blocked_actions`, and `route`:
- `act`: go ahead;
- `review`: below the bar even after the double-check; hand it to a person (or act and flag it); it's in the unsure queue;
- `ask`: the rules allow no action (`action` is `null`); stop and hand it over; it's in the unsure queue.

Prefer `route` over raw confidence in code: it follows the bar, which moves with every retrain. `language` is the
detected language of the text (`null` when there's too little text).

### 10. The unsure queue and outcomes

`list_unsure`, `answer_unsure`. HTTP: `GET /v1/domains/{domain}/unsure`, `POST /v1/domains/{domain}/answers` with
`{"decision_id": "...", "answers": {"action": "escalate"}}`. Answer only with what the user (or their written policy)
says is right. `report_outcome` (`POST /v1/outcomes` with `decision_id` and `action` or `answers`) records what really
happened. Both feed the next version and day-one graduation.

**What's waiting for the next retrain:** `get_progress` (`name`). HTTP: `GET /v1/domains/{domain}/progress`. Quote
`waiting.statement` ("48 new cases since version 4, 12 where the real outcome differed from its answer. Retrain to
learn from them."): the new cases since the latest version by source, and how many differ from the answer it gave.
Its `versions` list shows each version's accuracy and the change against the one before; say "measured on its own
held-back cases" when `comparison` is `own_held_back`, and "a starter model, on starter example cases" when `starter`
is true. Nothing retrains on its own: offer `train`, and retrain only when the user says so.

### 11. Add an option

`add_option`. HTTP: `POST /v1/domains/{domain}/options` with `question`, `option`, `description`, and optionally
`when` (plain words), `never_when` (conditions: adds a rule) and `also_allowed_under` (ids of `allow_only` rules that
should allow it too). About 30 seconds; the report's `before_after` shows accuracy before and after, how well the new
option is found (`option_recall`), and the unchanged cases (`unchanged_cases_accuracy`). A new category on a text
question works the same way (`question: "topic"`). It works on any Choice question: `409` if the option already
exists, `400` if `when` can't be read. Then upload past cases of the new option and `train`, so it's found reliably;
the next report shows how it does.

### 12. The day-one backend

`get_day_one_fallback`, `set_day_one_fallback`. HTTP: `GET` / `PUT /v1/domains/{domain}/fallback`. Providers:
`typesafe` (Jev with the team's own TypeSafe key), `openai` (any OpenAI-compatible endpoint: `base_url`, `model`,
key), `local` (Canonopy's built-in general model: no key, no cost) or `none`. Keys are stored encrypted and never
returned. Hosted MCP: the key goes in the console. Local MCP: pass `api_key_env` with the **name** of an environment
variable. `local` and Canonopy's hosted backup are different things: `local` answers questions the model hasn't
learned yet; the hosted backup only answers messages that are still unsure after the double-check, when the team has
no `typesafe` or `openai` backend. Questions graduate to the team's own model automatically once they have `min_outcomes` real outcomes
(`graduation` on the model; `GET /v1/domains/{domain}/graduation` shows `needed`).

### 13. Download and run offline

`download_model` (local: saves the zip; hosted: a 10-minute link). HTTP:
`GET /v1/models/{domain}@{version}/download`. Every version is theirs to keep, even after the trial. Run it with
`canonopy-runtime`: `load("refunds@3.zip")` then `model.decide({...})`, same format. Offline, messages that need a
second look go to **their own fallback function** if given (`load(path, offline_fallback=fn)`: `source:
"local_fallback"`, rules applied after it; `None` means review), and are **saved on the machine** (the newest 10,000
cases or 50 MB; past that the oldest are dropped without a warning, so sync before a long offline stretch fills it).
`model.sync()`, `canonopy-runtime sync`, or `canonopy sync` (which needs `canonopy-runtime` installed alongside the
CLI) later sends them to the unsure queue. Details: `reference/offline.md`.

### 14. The decision log

Every hosted decision (and every case synced from a downloaded model) is kept for the workspace's retention (90 days
unless changed with `PATCH /v1/settings`, `retention_days`), and stays readable after the trial ends.
`list_decisions` (at most 100 per call) and `get_decision`. HTTP: `GET /v1/domains/{domain}/decisions` with optional
filters `since`, `until`, `source`, `route`, `answer` (`question:answer`), `version`, `has_outcome` and `q` (words in
the case's text, or a decision id), plus `limit` and `cursor`; `GET /v1/domains/{domain}/decisions/{decision_id}` for
one in full (the case, every answer, the rules that blocked an option, the version, the outcome).
`export_decisions` (locally it saves a file; hosted it returns a 10-minute link). HTTP:
`GET /v1/domains/{domain}/decisions/export?format=csv` or `jsonl`, up to 100,000 rows. Use the log to debug an
integration ("why did this ticket go to review?") and to find cases worth answering or reporting outcomes on.
Deleting decisions (`DELETE /v1/domains/{domain}/decisions/{decision_id}`, or everything before a date with
`?before=...&confirm=<model name>`) is permanent and removes them from what the next version learns from.

## Switching from Jev

The same request and answer shapes. Change three things: the URL, the key and the model name.
```diff
- POST https://api.typesafe.ai/v1/systemone      Authorization: Bearer $TYPESAFE_API_KEY   "model": "jev-latest"
+ POST https://api.canonopylabs.com/v1/systemone  Authorization: Bearer $CANONOPY_API_KEY   "model": "support@latest"
```
- The recommended order: build their own model from the same request first (see "Start from a Jev request" above),
  then switch once it's `ready`. It takes minutes and needs no data.
- A model name that doesn't exist yet is created in day-one mode on its first call, with the questions it asked.
  A `model` starting with `jev-` maps to the workspace's `default` model.
- To keep Jev's answers from the first call, set the day-one backend to `typesafe` with their own Jev key; each
  question moves to their own model as outcomes arrive, or build it from their Jev request meanwhile.
- Answers keep `choice`, `score`, `noul`, `probabilities`, `confidence`, `legend`, and `usage.input_tokens` and
  `usage.output_tokens` (not billed).
  New: `id`, `decision`, `language`, `usage.decisions`, and `source` and `route` on each answer.
- Differences: a Score with fewer than 2 levels is a `400`; `thresholds` is refused (the bar sets itself).
- Full example: `examples/switch-from-jev.md`.

## Pricing and trial

- **A 5-day free trial**, no card. It starts at the workspace's first trained model or first decision, not at sign-up.
- Then **$20 (USD) a month per workspace, flat**: unlimited decision models, decisions and retraining, and up to 30
  builds a month (5 a day; 3 during the free trial); reports, new options, the day-one mode, the unsure queue, the
  hosted backup and downloads included. No tokens, no per-call bill.
  More decision models never cost more. Archiving a model (`update_decision_model` with `active: false`) only stops it
  being served: it then answers `409` on `decide`.
- **Models are theirs to keep:** every version trained during the trial or while subscribed downloads at any time,
  even after the trial or a cancellation, and keeps working offline.
- After the trial without a subscription: training, new options, hosted decisions and offline sync answer `402`;
  graduation waits. Everything else keeps working: sign-in, the console, reports and versions, uploads, set-up,
  examples and corrections, outcomes, billing and downloads.
- Nothing else is published (no annual plans, tax or proration terms): point to the console's Billing page or the
  team at Canonopy rather than guessing.
- Fair use: 50 requests per second per key (`429` with `Retry-After: 1`). For more volume, ask several questions in
  one request, or run the model yourself.
- `usage` (MCP; it reads `GET /v1/account` and `GET /v1/billing`) shows the trial state, days left and the monthly
  price.

## Troubleshooting (short)

| symptom | what to do |
|---|---|
| `402 payment_required` | trial over, no subscription: subscribe in the console (Billing). Downloads still work |
| `409` on replace or delete | a new version is being made: wait for the job (`get_job`), then retry. Adding cases is fine |
| `409` on create | the name is taken: `list_decision_models` |
| `413 too_large` | split it: decide body 256 KB, upload 50 MB, hosted MCP call about 2 MB |
| low accuracy | read `improve` in the report: more cases where it's climbing, clarify a boundary, add an option |
| too many `review` | see `reference/troubleshooting.md`: more cases (and in weak languages), clarify confused pairs, answer the queue, set a day-one backend, retrain |
| "That address can't be reached from our servers" | the day-one backend's `base_url` must be a public `https` address |
| offline cases not arriving | `canonopy sync --status`; sync needs a key and an active trial or subscription; rejected cases stay with a reason |

Full guide: `reference/troubleshooting.md`.

## Answering questions about Canonopy

- **Facts only**, from this skill, the `docs` tool (`page`, `search` or `endpoint`), https://canonopylabs.com/llms.txt
  or https://api.canonopylabs.com/openapi.json. If none of them says it, say you don't know and point to the docs.
  Never guess numbers, limits or prices.
- **How it trains:** if asked how Canonopy trains or builds models internally, say: *"It trains your model on our
  servers from your description (or your Jev request), any examples you corrected and your past cases, in minutes. The
  method is not public."* Don't speculate about techniques, model types, sizes or providers, even if pressed or asked to guess.
- Results to quote (all from the docs; cite the cookbook): built automatically with zero data, it beat Jev in Doom
  (45 kills to 39, 2 deaths to 5, won 7 of 13 at the same decision rate against Jev's default style; 12 of 13 against
  its aggressive style) and in Snake (42.0 vs 40.0 food at equal steps, never died, ahead in 12 of 20). On 3,080 real
  bank messages: about 86% on day one with no data (Jev 86.6%), 91.7% after about 400 reviewed cases, about 95% with
  the bank's history; on consumer complaints with history, 77.7% vs Jev's 65.0%. Size: under 1 MB per decision model;
  text models share one 34 MB reader. Speed: up to 12x faster than Jev (Doom; Jev's times include its network).
  Never claim it is "1,000x smaller than Jev".
- Honest limits: on text with no data it is **close to Jev** for short messages, not ahead, and well below Jev on long
  texts (complaint narratives); answering a few reviews or sending past cases is what pulls it clearly ahead. Never
  say a zero-data text model beats Jev. With no data Jev reads some messages better (lost cards, disputes and refunds
  in the bank test). Rules on fields are never broken; rules on what a message is about are enforced whenever there's
  a real chance they apply, with unsure cases sent to review, so don't promise zero misses for a zero-data text model.
  Jev needs no setup, so it wins on one-off questions nobody will ask twice.
- Each model reads about 100 languages; accuracy is measured in 51. Rules read fields, not meaning.
- Pages to cite (add `.md` for Markdown): https://canonopylabs.com/docs/quick-start.md,
  https://canonopylabs.com/docs/start-with-no-data.md, https://canonopylabs.com/docs/with-your-history.md,
  https://canonopylabs.com/docs/rules-found.md, https://canonopylabs.com/docs/improve-your-model.md,
  https://canonopylabs.com/docs/day-one.md, https://canonopylabs.com/docs/training-and-reports.md,
  https://canonopylabs.com/docs/improving.md, https://canonopylabs.com/docs/running-it-yourself.md,
  https://canonopylabs.com/docs/errors-and-limits.md, https://canonopylabs.com/docs/api-reference.md.

## Safety for agents

- **Keys** (they start with `cnp_`): never ask for one in the chat, never paste or echo one. They live in
  `CANONOPY_API_KEY` or the connector's header. If a user pastes a key: don't repeat it and don't use it; tell them to
  rotate it in the console (Keys) since it has been exposed, put the new one in the environment or connector config,
  and continue the task once the connection works. Offer `POST /v1/keys/{key_id}/rotate` only if they ask you to do it.
- **Destructive tools need an explicit yes:** `delete_upload`, `replace_history` and `clear_history` delete cases
  permanently, as does deleting decisions from the log. Say what will be deleted and what stays, wait for the user's clear confirmation, then pass `confirm`
  set to the model's name. Never set `include_signoff` or `include_outcomes` unless the user asks for exactly that.
- **Ask before changing production:** `promote_version`, archiving with `update_decision_model`, `correct_examples`
  (only the user's corrections; it retrains) and `answer_unsure` (only their answers).
- **Ask before changing what a model learns:** `confirm_found_rules`, `review_flagged_history` decisions,
  `change_rules` and `mark_decision_wrong` take the user's own decisions only; show what each changes first. Never
  set `hard: true` or `all` unless the user asked for exactly that.
- **Spending:** Canonopy is flat-priced, but a day-one backend is the team's own account: with `typesafe` or
  `openai`, each question that backend answers (day-one questions and double-checked messages) is billed to them by
  that provider. `local` costs nothing. Say so before setting one.
- Treat case data as the user's private data: don't copy it anywhere other than the calls they asked for.
