# Canonopy Decisions: the full docs > Canonopy Decisions: send us the JSON you send Jev. In minutes, get your own 1 MB model. For games and rule-based decisions, it beats Jev from day one. For text, it beats Jev clearly once it learns from your decisions. It is your own model for a decision you make over and over (route a ticket, approve a refund, flag a transaction, pick an NPC's move), asked in the same request and response format as TypeSafe's Jev (System One): Choice, Score and Noul questions. Send one real example of your state and your questions, exactly as you send them to Jev, to POST /v1/build (plus 5-20 more real states in `examples` and your rules in plain words if you like: no past decisions or answers needed; a few examples of what you send Jev make it better). Your model is ready in minutes, with nothing to do in between; review how it decides on about 20 examples any time (optional). Then switch the base URL and model: POST /v1/decide (alias /v1/systemone) keeps taking the same JSON. Your rules are enforced on every decision; the model is under 1 MB (text models share one 34 MB reader) and runs on our servers or yours, even offline. Built automatically with zero data it beat Jev in Doom (45 kills to 39, 2 deaths to 5, at the same decision rate) and Snake (42.0 to 40.0 food, never died). On text it is close to Jev on day one for short messages (bank: 84.6-86.8% vs 86.6%) and pulls clearly ahead as it learns from your decisions: 91.7% after about 400 reviewed cases, about 95% with the bank's history; on consumer complaints with history, 77.7% vs 65.0%. Past cases become the main examples, rules found in your history are proposed (optional to confirm), and every retrain (about a minute) is your call. Each decision model is called a domain in the API (/v1/domains). A 5-day free trial starts at your first model or first decision, no card; then $20 a month per workspace, flat, with unlimited decision models, decisions and retraining and up to 30 builds a month (5 a day, 3 during the trial). Your models are yours to keep. Index: https://canonopylabs.com/llms.txt. API description: https://api.canonopylabs.com/openapi.json. --- Source: https://canonopylabs.com/docs (Markdown: https://canonopylabs.com/docs.md) # Send us the JSON you send Jev In minutes, get your own 1 MB model: **a model for a decision you make over and over**, like routing a ticket, approving a refund, flagging a transaction or picking an NPC's next move. **No data needed.** You send the request you already send Jev, TypeSafe's System One API; we send back a model that answers the same questions in the same shape, with your rules enforced on every decision. It's under 1 MB, and it runs on our servers or yours, even offline. Then you use it: change the URL and the `model`, and keep sending the same JSON. For games and rule-based decisions, it beats Jev from day one. For text, it beats Jev clearly once it learns from your decisions: answer a few cases it was unsure about, or send your history, and retrain in a minute. - [Quick start](https://canonopylabs.com/docs/quick-start.md): Paste your Jev request, wait for ready, switch over. Curl, Python or the console. - [Start with no data](https://canonopylabs.com/docs/start-with-no-data.md): Everything about `POST /v1/build`: what it reads, the phases, the 95% quality check. - [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md): Your past cases become the main examples, and your model is measured on your own cases. - [Improve your model](https://canonopylabs.com/docs/improve-your-model.md): Advice on what helps most, targeted retrains with a before and after, rule changes that say what they change. - [Cookbooks](https://canonopylabs.com/docs/cookbooks/doom-zero-data.md): Doom and Snake built automatically with zero data from their Jev requests; bank support from day one to your history. > **For coding agents** > These docs are also plain Markdown. [llms.txt](https://canonopylabs.com/llms.txt) indexes every page, [llms-full.txt](https://canonopylabs.com/llms-full.txt) has all of them in one file, and adding `.md` to any docs URL gives that page. The API describes every endpoint and field at [api.canonopylabs.com/openapi.json](https://api.canonopylabs.com/openapi.json). Claude Code, Cursor and Codex can also use it directly through an MCP server: see [Use with coding agents](https://canonopylabs.com/docs/coding-agents.md). ## Against Jev The game players were built automatically by `POST /v1/build` from only what Jev is given (the state, the questions and the rules), with no past cases and no recorded play: | | Your model | Jev | |---|---|---| | **Doom**, zero data, 13 games at the same decision rate | **45 kills, 2 deaths**, won 7 of 13 | 39 kills, 5 deaths (default style) | | **Doom**, against Jev's aggressive style | **45 kills, 2 deaths**, won 12 of 13 | 21 kills, 11 deaths | | **Snake**, zero data, 20 games, food at equal steps | **42.0**, never died, ahead in 12 of 20 | 40.0, 1 crash (default strategy) | | **Bank support**, 3,080 real messages, day one with no data | 85.5% (10 automatic builds, 84.6–86.8%) | 86.6% few-shot, 83.9% zero-shot | | **Bank support**, after about 400 reviewed cases | **91.7%** | 86.6% few-shot | | **Bank support**, with the bank's history | **about 95%** | 86.6% few-shot | | **Consumer complaints** (CFPB), with history | **77.7%** | 65.0% few-shot, 62.9% zero-shot | On text, then: close to Jev on day one for short messages; give it past cases or answer a few reviews to pull clearly ahead. The numbers, how they were measured and every step are in the cookbooks: [Doom](https://canonopylabs.com/docs/cookbooks/doom-zero-data.md), [Snake](https://canonopylabs.com/docs/cookbooks/snake-zero-data.md) and [bank support](https://canonopylabs.com/docs/cookbooks/bank-zero-data.md). ## What you get - **Your own model, no data needed to start.** Built from the JSON you send Jev, in minutes. Each question has to agree with your instructions and rules at least 95% of the time on fresh situations before it serves. - **Rules you can count on.** Rules on your fields are enforced in code on every decision, and every answer says which rules blocked which actions. Rules on what a message is about are enforced whenever there's a real chance they apply, and unsure cases go to review. See [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md#rules). - **Calibrated confidence.** When it says 90%, it's right about 90% of the time. Route the unsure ones to a person. - **Small and fast.** Under 1 MB per decision model; text models share one 34 MB reader. Up to 12× faster than Jev in Doom (Jev's times include its network). Download it and run it on your own servers or offline. - **Better with your decisions.** Answer the cases it was unsure about, or send your history: your past cases become the main examples, [rules found in your history](https://canonopylabs.com/docs/rules-found.md) are proposed for you to confirm, and your model is measured on your own held-back cases. Every retrain is your call and takes about a minute. - **Flat price.** A 5-day free trial that starts when you first use it (your first model or first decision), no card. Then $20 a month per workspace, flat, with unlimited decision models, decisions and retraining and up to 30 builds a month. No tokens, no per-call bill. See [Errors and limits](https://canonopylabs.com/docs/errors-and-limits.md#build-limits). - **Yours to keep.** Every version can be downloaded at any time, even after the trial or if you cancel, and a downloaded model keeps working offline. - **Every decision on record.** A searchable decision log in the console and the API: what was decided, why, which rules applied, and the outcome. See [the decision log](https://canonopylabs.com/docs/decisions-and-rules.md#the-decision-log). Each decision you set up is a **decision model**. In the API it's called a *domain* (`/v1/domains`), and its name is what you pass as `model`, for example `refunds@latest`. ## Canonopy and Jev Jev is a general model that reads your question and answers it, with no setup. We answer the same questions with a model built for **your** decision. The trade: | | Jev | Canonopy Decisions | |---|---|---| | Setup | none | send the JSON you send Jev, ready in minutes; add history when you have it | | Accuracy on your decisions | good from the first call | games and rule-based decisions: beats Jev from day one (Doom, Snake); text: close to Jev on day one, clearly ahead once it learns from your decisions (91.7% after about 400 reviewed cases, about 95% with history, vs 86.6% in the bank test) | | Calibration error (bank test) | 0.044–0.063 | 0.006 with history | | Hard rules | yours to enforce in code | enforced on every decision, reported in the answer | | Where it runs | their cloud | our endpoint, or your own machines, even offline | | Price | per token | $20 a month per workspace, flat (unlimited decision models and decisions, up to 30 builds a month), after a 5-day free trial | > **Where Jev is better** > Jev needs no setup, so it wins on one-off questions nobody will ask twice. On text with no data, Jev is about as accurate on short messages and clearly better on long ones (complaint narratives), and it reads some messages more carefully: in the bank test it was better on messages about lost cards, refunds and disputes. Your rules on what a message is about are enforced whenever there's a real chance they apply, and unsure cases go to review; your history closes the gap. While a model is being built, or for a question it doesn't have yet, **your own Jev key** can answer in the same shape in [day-one mode](https://canonopylabs.com/docs/day-one.md). ## How it works 1. **Send the JSON you send Jev** to `POST /v1/build`, or paste it in the console: one real example of your `state`, your `questions` as they are, and your rules in plain words if you like. No past decisions or answers needed; a few examples of what you send Jev make it better (5-20 more real states in `examples`). See [Start with no data](https://canonopylabs.com/docs/start-with-no-data.md). 2. **Wait for `ready`.** Your model is built and checked in a few minutes, with nothing to do in between. Review how it decides any time (optional): about 20 situations with the answer it gives, and why. Correct any and retrain in a minute. 3. **Use it.** Change the base URL and the `model`, and keep sending the same JSON. 4. **Make it better when you like.** Add your history, or answer the cases it was unsure about, and retrain in about a minute: see [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md) and [Improve your model](https://canonopylabs.com/docs/improve-your-model.md). Prefer to describe the decision in plain words instead? The set-up agent does that: see [Training and reports](https://canonopylabs.com/docs/training-and-reports.md#set-up-the-domain). --- Source: https://canonopylabs.com/docs/quick-start (Markdown: https://canonopylabs.com/docs/quick-start.md) # Quick start Send us the JSON you send Jev. In minutes you get back your own model that answers the same questions, in the same shape, with your rules enforced. No past decisions or answers needed; a few examples of what you send Jev make it better. You need an API key from the [console](https://canonopylabs.com/console/keys). > Early access: the hosted API opens to teams on the [waitlist](https://canonopylabs.com/#waitlist) first. The base URL is `https://api.canonopylabs.com`. Prefer clicking? In the console, **New model**, then **Paste your Jev request**: paste it, paste a few more example states if you have them, add your rules, and when it's ready the page gives you the `model` to call. ## 1. Send the JSON you send Jev Take a request your code sends Jev today, as it is. Add 5-20 more real states in `examples` (states only, no answers: a quiet moment and a busy one, a small order and a large one), notes on what fields mean in `fields` if a name doesn't say it, your rules in plain words if you like, and a name. **curl** ```bash curl https://api.canonopylabs.com/v1/build \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "jev-latest", "state": { "order": { "amount": 42.5 }, "customer": { "tier": "pro", "fraud_flag": false } }, "questions": { "action": { "type": "choice", "instructions": "What should we do with this refund request?", "criteria": { "approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline" } }, "urgent": { "type": "noul", "instructions": "Does this need a person within the hour?" } }, "examples": [ { "order": { "amount": 820 }, "customer": { "tier": "free", "fraud_flag": false } }, { "order": { "amount": 12.99 }, "customer": { "tier": "pro", "fraud_flag": true } } ], "fields": { "order.amount": "the refund asked for, in USD (0-5,000)" }, "rules": "Never approve when amount is over 500. Always escalate when fraud_flag is true.", "name": "refunds" }' ``` **Python** ```python from canonopy import Client client = Client() # reads CANONOPY_API_KEY build = client.build( state={"order": {"amount": 42.5}, "customer": {"tier": "pro", "fraud_flag": False}}, questions={ "action": {"type": "choice", "instructions": "What should we do with this refund request?", "criteria": {"approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline"}}, "urgent": {"type": "noul", "instructions": "Does this need a person within the hour?"}, }, examples=[{"order": {"amount": 820}, "customer": {"tier": "free", "fraud_flag": False}}, {"order": {"amount": 12.99}, "customer": {"tier": "pro", "fraud_flag": True}}], fields={"order.amount": "the refund asked for, in USD (0-5,000)"}, rules="Never approve when amount is over 500. Always escalate when fraud_flag is true.", name="refunds", ) build = client.wait_for_build(build["build_id"]) # returns when the model is ready ``` **CLI** ```bash # jev-request.json: the body you send Jev today ({"model", "state", "questions"}) # more-states.jsonl: 5-20 more real states, one per line canonopy build jev-request.json --name refunds --examples more-states.jsonl \ --field "order.amount=the refund asked for, in USD (0-5,000)" \ --rules "Never approve when amount is over 500. Always escalate when fraud_flag is true." --wait ``` The answer comes at once. `model` is ignored, so a pasted request works as is. Check `understood` first: the fields read from your example (here `order_amount`, `customer_tier` and `customer_fraud_flag`), and the rules that will be enforced on every decision. With `fields`, `understood.field_notes` shows each note and the field it went to, and `understood.field_notes_not_matched` any note that matched no field. With a single state and no `examples`, `notes` reminds you that 5-20 real examples make a better model. ```json { "build_id": "bld_…", "domain": "refunds", "model": "refunds@latest", "status": "queued", "phase": "understanding", "understood": { "state": "object", "fields": [{ "name": "order_amount", "path": "order.amount", "type": "number" }, "…"], "decision_question": "action", "rules": ["Never approve when amount is over 500.", "Always escalate when fraud_flag is true."], "rules_not_understood": [] } } ``` ## 2. Wait for `ready` There is nothing to do in between: the build goes `queued` → `building` → `training` → `checking` → `ready`, usually in a few minutes. Follow it with `GET /v1/domains/refunds/build` (or `canonopy build status refunds`; `--wait` and `wait_for_build` return when it's ready). Each question has to agree with your instructions and rules at least **95%** of the time on fresh situations before it serves; `ready` means it passed. If one doesn't, the build is `needs_attention` and says where it's weakest and what to change. See [the quality check](https://canonopylabs.com/docs/start-with-no-data.md#the-quality-check-95-on-fresh-situations). ## 3. Review how it decides (optional) Your model is ready in minutes. Review how it decides any time (optional): about 20 situations, each with the answer your model gives and why. Nothing waits on them. ```bash curl https://api.canonopylabs.com/v1/domains/refunds/examples -H "Authorization: Bearer $CANONOPY_API_KEY" ``` If one is wrong, send only the ones you correct, with `"retrain": true` to retrain in the same call (about a minute): ```bash curl https://api.canonopylabs.com/v1/domains/refunds/signoff -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"examples": [{"id": "ex_02", "ok": false, "correct": {"action": "escalate"}}], "retrain": true}' ``` Or `canonopy examples refunds`, then `canonopy signoff refunds --correct ex_02:action=escalate --retrain`. The retrained model gives those answers. In the console, it's **Review examples** on the model's page. ## 4. Switch over and keep sending the same JSON Change two things in the code that calls Jev: the base URL and the model name. ```diff - POST https://api.typesafe.ai/v1/systemone "model": "jev-latest" + POST https://api.canonopylabs.com/v1/systemone "model": "refunds@latest" ``` `/v1/systemone` is an alias of `/v1/decide` with the same body. The request and answer shapes are Jev's, and `usage.input_tokens` is still there so client code that reads it doesn't break (it isn't billed). **curl** ```bash curl https://api.canonopylabs.com/v1/decide \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "refunds@latest", "state": { "order": { "amount": 820 }, "customer": { "tier": "pro", "fraud_flag": false } }, "questions": { "action": { "type": "choice", "instructions": "What should we do with this refund request?", "criteria": { "approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline" } }, "urgent": { "type": "noul", "instructions": "Does this need a person within the hour?" } } }' ``` **Python** ```python import os, requests res = requests.post( "https://api.canonopylabs.com/v1/decide", headers={"Authorization": f"Bearer {os.environ['CANONOPY_API_KEY']}"}, json={ "model": "refunds@latest", "state": {"order": {"amount": 820}, "customer": {"tier": "pro", "fraud_flag": False}}, "questions": { "action": {"type": "choice", "instructions": "What should we do with this refund request?", "criteria": {"approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline"}}, }, }, timeout=10, ) res.raise_for_status() decision = res.json()["decision"] print(decision["action"], decision["confidence"], decision["route"], decision["blocked_by_rules"]) ``` **JavaScript** ```ts const res = await fetch("https://api.canonopylabs.com/v1/decide", { method: "POST", headers: { Authorization: `Bearer ${process.env.CANONOPY_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ model: "refunds@latest", state: { order: { amount: 820 }, customer: { tier: "pro", fraud_flag: false } }, questions: { action: { type: "choice", instructions: "What should we do with this refund request?", criteria: { approve: "refund the customer", escalate: "a person decides", decline: "politely decline" } }, }, }), }); const { answers, decision } = await res.json(); ``` The answer: ```json { "id": "dec_4f1c2a9b0e7d6a51", "model": "refunds@1", "answers": { "action": { "type": "choice", "choice": "escalate", "confidence": 0.9705, "probabilities": { "approve": 0.0, "escalate": 0.9705, "decline": 0.0295 }, "source": "trained", "route": "act" }, "urgent": { "type": "noul", "noul": 0.09, "source": "trained", "route": "act" } }, "decision": { "question": "action", "action": "escalate", "confidence": 0.9705, "blocked_by_rules": ["never-approve-when-amount-is-over-500"], "blocked_actions": ["approve"], "route": "act" }, "language": "en", "usage": { "input_tokens": 118, "output_tokens": 0, "decisions": 1 } } ``` - `source` tells you who answered: your own model (`trained`), your day-one backend (`fallback`), or your rules alone (`rule`). - `route` on each answer, and `decision.route`, is `act` (automatic, at or above the model's safety bar) or `review` (held for a person); `decision.route` can also be `ask` (your rules allow no action). - `decision` is the final action with your rules applied: here the amount is over 500, so `approve` is blocked whatever the model thinks. See [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md). - Keep `id`: report what really happened with [`POST /v1/outcomes`](https://canonopylabs.com/docs/improving.md#report-outcomes), or [mark the decision wrong](https://canonopylabs.com/docs/improve-your-model.md#mark-a-decision-wrong). ## 5. Make it better, when you like Your model is ready with no data. It gets better as you use it, and every retrain is your call: - **Add your history.** Send your past cases in `history`, or upload them, and build again: your cases become the main examples, and your model is measured on your own held-back cases. See [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md). - **Confirm the rules found in your history.** Each one comes with how many past cases it covers and what confirming it changes. See [Rules found in your history](https://canonopylabs.com/docs/rules-found.md). - **Answer a few unsure cases, then retrain.** `GET /v1/domains/refunds/advice` says what would help most right now. A retrain takes about a minute and shows a before and after. See [Improve your model](https://canonopylabs.com/docs/improve-your-model.md). ## Just trying a call? The [Playground](https://canonopylabs.com/console/playground) sends a state to any of your models and shows each answer with its source, the decision and the rules that fired. Every run has a share link. A `model` that doesn't exist yet is created on its first call in [day-one mode](https://canonopylabs.com/docs/day-one.md), with the questions that call asked: until it has a model of its own, your day-one backend (your own Jev key, an OpenAI-compatible model or our built-in one) answers in the same shape (`"source": "fallback"`), and a day-one answer below 0.8 confidence is `review`. A `model` starting with `jev-` maps to your `default` model. ## Other ways in The API is all you need: every way in below is optional. The one package you'd install is `canonopy-runtime`, and only to run a downloaded model yourself. - **Coding agents:** connect Claude Code, Cursor or Codex to the hosted MCP connector, nothing to install: `claude mcp add --transport http canonopy https://api.canonopylabs.com/mcp --header "Authorization: Bearer cnp_…"`. Then ask it to *"build a Canonopy model from the Jev request in `src/refunds.py`"*. See [Use with coding agents](https://canonopylabs.com/docs/coding-agents.md). - **CLI:** the `canonopy` command for builds, examples, versions and more: `pip install https://canonopylabs.com/dl/canonopy_cli-0.6.2-py3-none-any.whl`. See the [CLI reference](https://canonopylabs.com/docs/cli-reference.md). - **Python SDK:** mirrors Jev's `Choice`, `Score` and `Noul` (`from canonopy import Client`), and `@canonopy.fn` in place of `@jev.fn`: `pip install https://canonopylabs.com/dl/canonopy_decisions-0.4.0-py3-none-any.whl`. See [Start with no data](https://canonopylabs.com/docs/start-with-no-data.md#from-a-python-function). - **Run it yourself:** download a version and run it with `canonopy-runtime`. See [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md). --- Source: https://canonopylabs.com/docs/start-with-no-data (Markdown: https://canonopylabs.com/docs/start-with-no-data.md) # Start with no data Send us the same JSON you send Jev. Get your own model back in minutes. No past decisions or answers needed; a few examples of what you send Jev make it better: one real example of your `state`, 5-20 more in `examples`, your `questions` exactly as they are, and, if you like, notes on your fields and your rules in plain words. **`POST /v1/build`** Alias: `POST /v1/systemone/build`, same body. A pasted Jev request works as is: `model` and anything else Jev takes are ignored. **curl** ```bash curl https://api.canonopylabs.com/v1/build \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "state": { "order": { "amount": 42.5 }, "customer": { "tier": "pro", "fraud_flag": false } }, "questions": { "action": { "type": "choice", "instructions": "What should we do with this refund request?", "criteria": { "approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline" } }, "urgent": { "type": "noul", "instructions": "Does this need a person within the hour?" } }, "examples": [ { "order": { "amount": 820 }, "customer": { "tier": "free", "fraud_flag": false } }, { "order": { "amount": 12.99 }, "customer": { "tier": "pro", "fraud_flag": true } } ], "fields": { "order.amount": "the refund asked for, in USD (0-5,000)" }, "rules": "Never approve when amount is over 500. Always escalate when fraud_flag is true.", "name": "refunds", "description": "Refund requests" }' ``` **Python** ```python from canonopy import Client client = Client() # reads CANONOPY_API_KEY build = client.build( state={"order": {"amount": 42.5}, "customer": {"tier": "pro", "fraud_flag": False}}, questions={ "action": {"type": "choice", "instructions": "What should we do with this refund request?", "criteria": {"approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline"}}, "urgent": {"type": "noul", "instructions": "Does this need a person within the hour?"}, }, examples=[{"order": {"amount": 820}, "customer": {"tier": "free", "fraud_flag": False}}, {"order": {"amount": 12.99}, "customer": {"tier": "pro", "fraud_flag": True}}], fields={"order.amount": "the refund asked for, in USD (0-5,000)"}, rules="Never approve when amount is over 500. Always escalate when fraud_flag is true.", name="refunds", description="Refund requests", ) print(build["build_id"], build["status"], build["message"]) ``` **CLI** ```bash # jev-request.json: the body you send Jev today ({"model", "state", "questions"}) # more-states.jsonl: 5-20 more real states, one per line (or a JSON list) canonopy build jev-request.json --name refunds --examples more-states.jsonl \ --field "order.amount=the refund asked for, in USD (0-5,000)" \ --rules "Never approve when amount is over 500. Always escalate when fraud_flag is true." canonopy build status refunds # where it is; the build id works too ``` - `state` (string | object, required): ONE real example: a text, or the JSON object your program sends. Every value is read as a field with a type and a `path`: numbers, `true`/`false` (yes/no), short codes (categories) and longer writing (text). - `examples` ((string | object)[]): Optional, recommended: 5-20 more real states, each like `state` (up to 50; states only, no answers). Pick different ones: a quiet moment and a busy one, a small order and a large one. Your model's situations are made from all of them and checked against them, and your fields' ranges and units are read from them. With one state only, `notes` says so. - `fields` (object): Optional: what fields mean, keyed by a field's `path` (as in `understood.fields`, e.g. `player.angle`) or its name. A note is a sentence, e.g. `"turn angle in engine units (0-4,294,967,295)"`, or `{"meaning", "unit", "range": [lowest, highest]}`. Your model's situations follow them. - `questions` (object, required): Your questions, exactly as you send them to Jev: Choice, Noul and Score, with instructions and criteria. They're stored as they are, so you keep sending the same JSON. - `rules` (string): Your rules in plain words, e.g. "Never approve when amount is over 500. Always escalate when fraud_flag is true." - `name` (string): The model's name, what you'll pass as `model`. Default: one made from `description` or your question ids. Building again under the same name makes the next version of that model. - `description` (string): What it decides, in a sentence. - `history` (object[]): Optional: your past cases with their answers. See [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md). ## What it reads from your example The answer comes at once, with what was read in `understood`. Check it before anything else: - `fields`: every field your model reads, with its `path` and its type (`number`, `yes/no`, `category` or `text`). Lists of objects are read at their first 8 positions (`[0]` to `[7]`); lists of plain values one position per item. - `not_read`: what isn't read, and why. Identifiers (`id`, `*_id`, uuids, emails, phone numbers, links, timestamps), empty values and empty lists are left out on purpose. - `field_notes`: each note from `fields`, with the `path` it went to and its `meaning`, `unit` and `range`. `field_notes_not_matched`: a note that matched no field (`key`), and `why`. - `questions`: your questions, as stored. `decision_question`: the Choice question your hard rules apply to. - `rules`: the rules read and enforced on every decision. `rules_not_understood`: what couldn't be read as a rule that blocks answers, each with the reason. Your model still follows it in its answers; it just isn't enforced as a hard rule. ## Rules in plain words Write them the way you'd say them, one per sentence: "Never approve when amount is over 500." "Always escalate when fraud_flag is true." "When tier is free, only decline or escalate." Name a field by its path (`order.amount`), its name (`order_amount`) or, when it's unique, its last key (`amount`). Rules that block answers of your decision question are **enforced on every decision**, in code, whatever the model thinks, and reported in each answer. A rule on another question's answer ("Always block when topic is lost_card") is enforced whenever there's a real chance it applies. See [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md#rules). ## The phases Follow the build with `GET /v1/build/{build_id}`, or the model's latest build with `GET /v1/domains/{domain}/build`. The console shows the same progress. **`GET /v1/build/{build_id}`** **`GET /v1/domains/{domain}/build`** | phase | what happens | |---|---| | `understanding` | reading your example and your questions | | `preparing` | preparing examples of situations for your questions, and about 20 examples of how your model decides | | `training` | your model is being built, on our servers | | `checking` | your model against your instructions and rules, on fresh situations | `status` goes `queued` → `building` → `training` → `checking` → `ready`, or `needs_attention` (built, not serving yet: see below), or `failed` (`errors` says why, in plain words). Nothing waits for you in between. `message` and `next_step` always say where it is and what to do. ```json { "build_id": "bld_…", "domain": "refunds", "model": "refunds@latest", "status": "training", "phase": "training", "message": "Your model is being built.", "next_step": "Your model is being built. Check this build again in a minute or two.", "understood": { "state": "object", "fields": [{ "name": "order_amount", "path": "order.amount", "type": "number" }, "…"], "decision_question": "action", "rules": ["Never approve when amount is over 500.", "Always escalate when fraud_flag is true."], "rules_not_understood": [] }, "examples": [], "quality": null, "stats": { "seconds": 94.2, "situations": 30000, "examples": 20 } } ``` ## The quality check: 95% on fresh situations Once trained, your model is checked against your instructions and rules on fresh situations it never learned from. Each question needs to agree at least **95%** of the time (`quality.bar`) before the model serves. ```json "quality": { "bar": 0.95, "passed": true, "situations": 4000, "questions": [{ "question": "action", "agreement": 0.978, "passed": true, "cases": 4000, "weakest": [{ "where": "when the right answer is 'escalate'", "agreement": 0.93, "cases": 240 }] }], "statement": "…" } ``` - **`ready`:** it serves as `refunds@latest`. - **`needs_attention`:** the version is kept but doesn't serve. Each question under the bar has `advice` in plain words: where it's weakest ("… weakest when the right answer is 'escalate' (72%) and when amount is under 150 (80%) …") and what to change. Clarify your instructions or rules and build again under the same name, or [retrain it](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain). - `weakest` lists where each question matches least, even when it passes: what a retrain sharpens. `stats` has how long it took, the number of situations your model learned from, and once trained the download's size (`download_bytes`) and the unzipped model's (`model_file_bytes`). The version's report has the same check, with its questions `measured_on: "situations"`. Once you have enough answered cases of your own (about 200: answer the review queue or report outcomes), a new version is also compared with the serving one on your own held-back cases, and it serves only if it isn't worse there (`quality.measured_on: "your_cases"`, with the numbers in `quality.compared`). ## Review how it decides (optional) Your model is ready in minutes. Review how it decides any time (optional): about 20 situations, each with the answer your model gives and why. They're at `GET /v1/domains/{domain}/examples`, and in the build's `examples` once it's `ready` or `needs_attention`. Nothing waits on them. To correct one, send only the ones you correct. With `"retrain": true` the retrain starts in the same call: **`POST /v1/domains/{domain}/signoff`** ```json { "examples": [ { "id": "ex_02", "ok": false, "correct": { "action": "escalate" }, "note": "pro accounts under 30 days go to a person" } ], "retrain": true } ``` The answer has a `message` and the retrain's `build_id`. Your corrections also go to your model's instructions, so the retrained model gives those answers. Without `retrain`, they apply at your next retrain (`POST /v1/domains/{domain}/train`). The review is kept as a record. In the CLI: `canonopy examples refunds`, then `canonopy signoff refunds --correct ex_02:action=escalate --retrain`. ## Switch over and keep sending the same JSON When it's `ready`, change two things in the code that calls Jev today: the base URL and the model name. The request and the answer keep their shape. ```diff - POST https://api.typesafe.ai/v1/systemone "model": "jev-latest" + POST https://api.canonopylabs.com/v1/systemone "model": "refunds@latest" ``` `/v1/systemone` is an alias of `/v1/decide`. Every answer also says who answered (`source`) and whether to act on it (`route`), and your rules are applied in `decision`. See [Quick start](https://canonopylabs.com/docs/quick-start.md#4-switch-over-and-keep-sending-the-same-json) and [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md). ## From a Python function: `@canonopy.fn` If you describe the decision as a typed Python function, put `@canonopy.fn` on it (swap `@jev.fn` for `@canonopy.fn`). It needs `canonopy-decisions` 0.2.0 or later and pydantic 2. Only the function's definition is read; its body is never run. - **The parameters are your state's fields:** `str` is text, `int` and `float` numbers (`Field(ge=…, le=…)` bounds them), `bool` yes/no, a `Literal` or an `Enum` a category, a pydantic model a nested object. - **The return type, a pydantic model, is your questions:** a `Literal` or an `Enum` is a Choice (option descriptions in `Field(json_schema_extra={"options": {...}})`), a `bool` a Noul, a bounded `int` a Score. A field's `description` is its question's instructions. - **The docstring:** its first paragraph is the description, and the rest are your rules in plain words. ```python from typing import Literal from pydantic import BaseModel, Field import canonopy class Refund(BaseModel): action: Literal["approve", "escalate", "decline"] = Field( description="What should we do with this refund request?", json_schema_extra={"options": {"approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline"}}) urgent: bool = Field(description="Does this need a person within the hour?") @canonopy.fn(name="refunds") def refund(message: str, amount: float, tier: Literal["free", "pro", "enterprise"], fraud_flag: bool) -> Refund: """Refund requests from our shop. Never approve when amount is over 500. Always escalate when fraud_flag is true. """ refund.request # the JSON it sends to POST /v1/build: name, state, questions, rules, description build = refund.build(wait=True) # waits until the model is ready (or needs attention, or failed) answer = refund("Charged twice, please refund", amount=820, tier="pro", fraud_flag=False) answer.action # "escalate": a Refund, with your rules applied ``` `@canonopy.fn(name=None, *, rules=None, example=None, description=None, decision=None, client=None, examples=None, fields=None)`: - `name`: the model's name (default: the function's); `rules`: more rules in plain words, before the docstring's; `description`: instead of the docstring's first paragraph. - `example`: the example `state` to build from (default: one made from the parameter types and their defaults). `examples`: 5-20 more real states; `fields`: notes on what fields mean. - `decision`: the question that gets the final decision's action (default: the one your rules apply to); `client`: the `Client` to use. `.build(history=[...])` sends your past cases too. `.bind(model="refunds@3")` points the function at a fixed version, or a model you already have. Calling the function decides with `refunds@latest` and returns the typed answer: the decision question gets the action with your rules applied, a Noul `True` or `False`, a Score its level. ## From a coding agent The MCP tools (the hosted connector and `canonopy mcp`, see [Use with coding agents](https://canonopylabs.com/docs/coding-agents.md)): - `build_model`: `state`, `questions`, and optionally `examples` (5-20 more real states), `fields`, `rules`, `name`, `description` and `history` rows (the local server also takes `history_file`). - `get_build`: a `build_id`, or a `model` for its latest build. - `get_examples`, `correct_examples`: optional, once it's ready: how your model decides on about 20 situations, and corrections (with a retrain). - `decide`: the same JSON, with `model` set to the new name. A prompt that works: *"Take the Jev request in `src/refunds.py`, build a Canonopy model from it with our refund rules, and switch the code over once it's ready."* Then, if you like: *"Show me how it decides on the examples."* ## In the console **New model**, then **Paste your Jev request**: paste the JSON, paste a few more examples (one state per line) and notes on fields if you like, add your rules, and follow the build. Once it's ready, the page gives the `model` to call, and **Review examples** shows how it decides (optional). ## Good to know - **In minutes:** preparing the examples takes a few minutes, training about a minute. Your model is small enough to [download and run yourself](https://canonopylabs.com/docs/running-it-yourself.md). - **Limits:** 5 builds per workspace per UTC day, 3 during the free trial, 30 a month with a subscription (`429`, or `402` when the trial's builds are used up). A build is a new model, a new version, or a rule change; retrains don't count. A build counts as training for your plan: after the free trial it needs a subscription (`402`). You're never charged per build. See [Build limits](https://canonopylabs.com/docs/errors-and-limits.md#build-limits). - **Errors:** `400` when the example or the questions can't be read (an empty object, a Choice with one option), `409` when the name belongs to a model not made this way, or a build of it is still in progress. - **Past cases later:** your uploads, outcomes and unsure-queue answers improve the same model. See [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md) and [Improve your model](https://canonopylabs.com/docs/improve-your-model.md). --- Source: https://canonopylabs.com/docs/build-with-your-coding-agent (Markdown: https://canonopylabs.com/docs/build-with-your-coding-agent.md) # Build with your coding agent The fastest way to your own model: your coding agent talks to ours. It already knows your code; we know how to make a fast model from it. You stay in charge of every decision. 1. **Connect it.** Add the MCP connector once (see [Use with coding agents](https://canonopylabs.com/docs/coding-agents.md)): ```bash claude mcp add --transport http canonopy https://api.canonopylabs.com/mcp \ --header "Authorization: Bearer $CANONOPY_API_KEY" ``` 2. **Ask for a model.** "Build a Canonopy model for our refund decisions." That's all it needs to hear. ## JSON first Nothing is built until your agent has the same JSON you send Jev: one real `state` and your `questions`. If you don't hand it over, `build_model` builds nothing and tells your agent how to get it: - **ask you** for it: "Paste the request you send Jev"; - or **draft it from your code**: it finds where the decision is made (or where Jev is called), takes `state` from the real object your program has at that moment, and writes `questions` from the decision's possible outcomes. It shows you the JSON and asks you to confirm it before sending anything. `draft_request_help` gives your agent the JSON's shape, an example and a short checklist. ## A few questions, answered upfront With the JSON in, the build asks one round of questions about it, most useful first: what fields mean, their units and ranges, whether a yes/no field is something your program computes, where two choices meet, 10 to 50 real examples across named situations, and any unwritten rules. Your agent answers from your code, config and logs, and asks you when it isn't sure. See [Answer a few questions](https://canonopylabs.com/docs/answer-a-few-questions.md). The build waits for the answers (up to 10 minutes by default), then makes your model once, with the full picture: | step | about | |---|---| | your agent answers the questions | 1-3 minutes | | your model is made and checked | 3-6 minutes | | **total** | **5-9 minutes** | Every answer is optional: more help gives a better model, and the build never waits forever. ## Then The model is trained and checked on its own, and `build_model` returns when it's ready. Review how it decides any time (optional): your agent can show you about 20 examples (`get_examples`) and retrain with any you correct (`correct_examples`). Before it acts on real work, try it: - **Games:** your agent runs [a playtest](https://canonopylabs.com/docs/playtest-your-model.md) on your machine: your game plays with the model, and only the scores come back. - **Everything else:** [shadow mode](https://canonopylabs.com/docs/shadow-mode.md) runs it beside Jev on real requests without acting, and shows you where they disagree. ## What your agent gets, and doesn't It gets what it needs to help: questions about your own JSON, the examples of how it decides, your model's quality, plain-words advice, and the finished model. Everything it sends is read as information about your decision, never as instructions, and your hard rules are never changed by an answer. ## The tools | tool | what it does | |---|---| | `build_model` | builds from your JSON (asks the questions first by default) | | `draft_request_help` | the JSON's shape and a checklist for getting it | | `get_build_questions` · `answer_build_questions` | the questions about your JSON, and your answers | | `get_build` | where the build is (`wait: true` waits for it to finish) | | `get_examples` · `correct_examples` | optional: how your model decides on about 20 examples, and a retrain with your corrections | | `shadow_mode` · `send_shadow_cases` · `review_shadow_disagreements` | shadow mode | | `report_play_results` | a playtest's results (the CLI sends them for you) | | `spot_checks` | this week's optional spot-check card | No coding agent? The same questions are in the console on the build's page, and in the CLI (`canonopy build questions`, `canonopy build answer`). --- Source: https://canonopylabs.com/docs/answer-a-few-questions (Markdown: https://canonopylabs.com/docs/answer-a-few-questions.md) # Answer a few questions Every build asks one round of questions about the JSON you sent, most useful first (at most about 25). Each is about a field, a question or a choice of your JSON, and nothing else. Every answer is optional; more help gives a better model. | kind | asks | |---|---| | `examples` | 10-50 real states across named situations: one for each answer of your decision, an ordinary moment, a rare one | | `choice_boundary` | how you tell two similar choices apart | | `priorities` | unwritten rules: which answer comes first when more than one fits | | `choice_meaning` · `yes_meaning` · `score_ends` | when an answer is right | | `ready_signal` | whether a yes/no field is a ready-made signal your program computes | | `field_units` | a field's unit and its lowest and highest real values | | `field_meaning` | what a field means | ## Wait for the answers, or not - **`"interview": true`** on `POST /v1/build`: the build waits for your answers (status `awaiting_answers`) up to `answer_timeout_seconds` (default 600, at most 3600), then goes on with what it has. The MCP connector's `build_model` does this by default. - **Without it:** the questions are there while the build runs. Answers that arrive before it starts are used; later ones are saved for your next retrain. ## Read them **`GET /v1/build/{build_id}/questions`** ```json { "build_id": "bld_…", "round": 1, "status": "open", "answer_by": "2026-09-30T14:10:00Z", "questions": [ { "id": "q01", "kind": "examples", "priority": 1, "text": "Send 10 real states your program sends, across these situations: one where `action` is `approve`; …", "refs": { "question": "action", "choices": ["approve", "escalate", "decline"], "fields": [] }, "answer_with": { "examples": "a list of {\"situation\", \"state\", \"answers\"}" } }, { "id": "q02", "kind": "choice_boundary", "priority": 2, "text": "In `action`, how do you tell `approve` from `escalate`? What tips a case from one to the other?", "refs": { "question": "action", "choices": ["approve", "escalate"], "fields": [] }, "answer_with": { "text": "a sentence or two (up to 600 characters)" } } ] } ``` ## Answer them **`POST /v1/build/{build_id}/answers`** ```json { "answers": [ { "id": "q01", "examples": [ { "situation": "a large refund", "state": { "order": { "amount": 820 }, "customer": { "tier": "free" } }, "answers": { "action": "escalate" } } ] }, { "id": "q02", "text": "Escalate when a person should look: unusual amounts or an upset customer." }, { "id": "q05", "unit": "dollars", "range": [0, 5000] }, { "id": "q06", "ready_made": true, "text": "our fraud service sets it" }, { "id": "q07", "skip": true } ], "done": true } ``` - `answers` (object[], required): One per question: its `id` and the keys its `answer_with` lists: `text` (up to 600 characters), `unit` and `range` ([lowest, highest]), `ready_made` (true or false), `examples` (up to 50 real states like the one you sent, each with a short `situation` and, if you know it, the `answers` you'd expect), or `skip`. - `done` (boolean): Default `true`: that's all, the build goes on now. `false`: more answers follow, and it keeps waiting up to its timeout. - `round` (integer): The round you answer (default: the latest). The answer says what was recorded (`accepted`), what wasn't and why (`rejected`), and where your answers go (`used`: `this_build` or `next_retrain`). Your answers are read as information about your decision: they add notes on your fields, real examples your model's situations are made from, and what you said about your choices. They never change your hard rules. ## A short follow-up, only when it helps After the check, a build that asked its questions may ask a few more (at most 6), and only when it found a real gap: two choices it mixes up, or a question below the bar. It never holds your model up: answer them and retrain for a better version (`POST /v1/domains/{domain}/train`). The advice points them out. ## In the console and the CLI The build's page has a form for the questions. In the CLI: ```bash canonopy build jev-request.json --interview # the build waits for your answers canonopy build questions BUILD_ID canonopy build answer BUILD_ID # asks each question on the terminal (Enter skips) canonopy build answer BUILD_ID --file answers.json ``` --- Source: https://canonopylabs.com/docs/what-you-need-to-bring (Markdown: https://canonopylabs.com/docs/what-you-need-to-bring.md) # What you need to bring **Nothing but your Jev request.** One real example of the `state` your program sends, your `questions` exactly as you send them to Jev, and, if you like, your rules in plain words. No past decisions or answers needed; a few examples of what you send Jev make it better: 5-20 more real states in `examples`. That's enough for your own model, in minutes: see [Start with no data](https://canonopylabs.com/docs/start-with-no-data.md). **Your decisions make it better.** Past cases with the answers they got, answers to cases your model was unsure about, reported outcomes, or recorded play: whatever you have, whenever you have it. See [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md). ## What that looks like | Your decisions | What the model reads | Examples | To start | Makes it better | |---|---|---|---|---| | On human language | free text | support tickets, claims, complaints, moderation | your Jev request | past cases and how each was handled | | On data | numbers and fields | fraud flags, eligibility, reorders, alert triage | your Jev request | reported outcomes and decisions marked wrong | | Step by step | a changing state | game opponents, NPCs, agents in a simulator | your game's Jev request | recorded play | | Mixed | fields plus some text | a bank case with account details and a message | your Jev request | past cases, for the text part | ## Decisions on text Tickets, emails, chats, claims, reviews. Close to Jev on day one for short messages; give it past cases or answer a few reviews to pull clearly ahead. On 3,080 real bank messages: | | Accuracy | |---|---| | Your model, day one with no data (10 automatic builds) | 85.5% (84.6–86.8%) | | Your model, after about 400 reviewed cases | **91.7%** | | Your model, with the bank's history | **about 95%** (94.9–96.0%) | | Jev, few-shot | 86.6% | | Jev, zero-shot | 83.9% | On long texts, like complaint narratives, a model built with no data is well below Jev; there, start from your past cases. With history, on real consumer complaints (CFPB): **77.7%** against Jev's 65.0% few-shot. - **The answers you already record are enough.** The outcome or category you already keep for each past case is what your model learns from. A column named like a question holds how the case went. - **Rules on what a message is about** ("Always block-card when topic is lost_stolen") are enforced whenever there's a real chance they apply, and unsure cases go to review. With your history, your model reads those messages better too. See [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md#rules-on-what-a-message-is-about). - **Answering the unsure cases helps most.** Each answer to a case your model was least sure about teaches the next retrain the most: see [Improve your model](https://canonopylabs.com/docs/improve-your-model.md#advice). - It reads about 100 languages; accuracy is measured in 51. See [Languages](https://canonopylabs.com/docs/languages.md). The numbers and every step: [Bank support, from day one](https://canonopylabs.com/docs/cookbooks/bank-zero-data.md) and [with history](https://canonopylabs.com/docs/cookbooks/bank-support.md). ## Decisions on data and state Transactions, account fields, sensor readings, schedules. Your Jev request is all there is to bring. Numbers go into the model as numbers, so it can tell £4,999 from £5,001, and your rules on fields are enforced on every decision, whatever the model thinks. ## Games A move in a game is a decision on data. Send the request your game sends Jev: its state, with each move's facts if you have them, and a Choice question over the moves. Call your model like Jev, or download it and run it inside the game loop. Built automatically with zero data from their Jev requests, a Doom player beat Jev at the same decision rate (45 kills to 39, 2 deaths to 5, won 7 of 13 against its default style and 12 of 13 against its aggressive one), and a Snake player out-ate Jev's (42.0 to 40.0 food at equal steps, never died). See the [Doom](https://canonopylabs.com/docs/cookbooks/doom-zero-data.md) and [Snake](https://canonopylabs.com/docs/cookbooks/snake-zero-data.md) cookbooks. Recorded play makes it better: moments from a scripted bot, a designer or your best players, each with the right move. See the history cookbooks for [Doom](https://canonopylabs.com/docs/cookbooks/doom.md) and [Snake](https://canonopylabs.com/docs/cookbooks/snake.md). ## The shape of past cases When you have them: a CSV with your state fields and one column per question you have answers for: ```text message,amount,tier,fraud_flag,action "Refund my order please",40,pro,0,approve "I was charged 900 twice",900,enterprise,0,escalate ``` Or JSONL, one case per line, with the state exactly as your program sends it: `{"state": {...}, "answers": {"action": "approve"}}`. Send them in `history` with your build, or upload them: see [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md#send-your-past-cases). --- Source: https://canonopylabs.com/docs/questions (Markdown: https://canonopylabs.com/docs/questions.md) # Questions: Choice, Score and Noul A question is one narrow judgment. You ask as many as you like in one request; each gets its own answer. The format is Jev's, field for field. ```json { "model": "bank-support@latest", "state": { "message": "…", "amount": 86.5 }, "questions": { "topic": { "type": "choice", "instructions": "What is the customer writing about?", "criteria": { "lost_stolen": "lost, stolen or compromised card", "refund": "wants a refund" } }, "urgency": { "type": "score", "instructions": "How urgent is it?", "criteria": ["can wait", "today", "right now"] }, "needs_human": { "type": "noul", "instructions": "Does this need a person?", "criteria": { "true": "a person must handle it", "false": "can be handled automatically" } } } } ``` Leave out `questions` to ask every question your domain knows. `instructions` and every criterion can be text, an object or an array. ## Choice One of up to 255 named options. ```json { "type": "choice", "choice": "lost_stolen", "confidence": 0.93, "probabilities": { "lost_stolen": 0.93, "unrecognised": 0.06, "refund": 0.004 }, "source": "trained" } ``` - `choice` is the most likely option; `probabilities` add up to 1. - `confidence` is the probability of `choice`. - Write contrastive criteria: say what an option is for, and what it's **not** for. It helps your day-one backend, and it's how options are described to your model too. ## Score A position on 2 to 10 ordered, described levels. ```json { "type": "score", "score": 1.62, "confidence": 0.58, "legend": { "0": "can wait", "1": "today", "2": "right now" }, "probabilities": { "0": 0.02, "1": 0.34, "2": 0.64 }, "source": "fallback" } ``` - `score` is the probability-weighted level: the sum of each level times its probability. - `confidence` is the largest single probability. - `legend` echoes your level descriptions. *Interactive on the web page (https://canonopylabs.com/docs/questions): how a Score's probabilities give its score.* ## Noul The probability that a yes-or-no statement is true, from 0 to 1. ```json { "type": "noul", "noul": 0.31, "source": "fallback" } ``` A Noul has no separate `confidence`: the number is the confidence. Treat 0.5 as "no idea" and set your own cut-offs in code. ## Where each answer comes from Every answer carries `source`: | `source` | Who answered | |---|---| | `trained` | your own model, learned from your cases | | `fallback` | your day-one backend, until the question graduates ([Day one](https://canonopylabs.com/docs/day-one.md)) | | `rule` | your rules left only one option | ## Good questions - **Atomic.** One judgment per question. "Is this a refund request?" and "How urgent is it?" beat one question that tries to do both. - **Stable.** Questions you ask on every case are the ones worth a model of their own. One-off questions are fine too: they're answered by your fallback. - **Combine in code, or declare a decision.** Weights, cut-offs and routing live in your code, or in the domain's [decision](https://canonopylabs.com/docs/decisions-and-rules.md). --- Source: https://canonopylabs.com/docs/decisions-and-rules (Markdown: https://canonopylabs.com/docs/decisions-and-rules.md) # Decisions and rules Most domains have one question whose options are the **actions**: route, refund, escalate, block. That's the domain's `decision_question`. When it's asked, the answer carries a `decision` block with your rules applied. ```json "decision": { "question": "action", "action": "escalate-senior", "confidence": 0.91, "blocked_by_rules": ["refund-over-250-needs-a-human"], "blocked_actions": ["refund"], "route": "act" } ``` | Field | Meaning | |---|---| | `question` | the question whose options are the actions | | `action` | the final action with your rules applied; `null` if your rules block every action | | `confidence` | how sure it is of that action, after the rules | | `blocked_by_rules` | the rules that blocked at least one action for this case | | `blocked_actions` | the actions they blocked | | `rule_chances` | only when a rule on what the message is about applied: each such rule and the chance it applies here (see below) | | `route` | `act`, `review` or `ask`, from the model's safety bar | The decision question's answer comes back with your rules applied too, so blocked options are at 0 and `answers[decision.question].choice` equals `decision.action`. ## Rules A rule says which actions are never allowed, or the only ones allowed, when its conditions hold. Rules read your **state fields**, or **another question's answer** (what a message is about; see below). ```json { "id": "over-500-needs-a-person", "description": "Never approve more than $500 without a person", "when": [{ "field": "amount", "op": "gt", "value": 500 }], "block": ["approve"] } { "id": "fraud-goes-to-a-person", "description": "Fraud-flagged accounts always go to a person", "when": [{ "field": "fraud_flag", "op": "is_true" }], "allow_only": ["escalate"] } ``` - **Operators:** `eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `in`, `not_in`, `is_true`, `is_false`, `is_missing`, `contains`. - **Combine** with `{"any": [...]}` and `{"all": [...]}`. Every condition in `when` must hold. - **Enforced on every decision,** in code, whatever the model thinks. Training reports count how often each rule applied and changed a decision; violations are always 0. You rarely write these by hand: describe them to the set-up agent in plain words ("never approve over 500") and it adds them. You can also `PATCH /v1/domains/{domain}` with a `rules` list. ### Rules on what a message is about A rule can read another question's answer instead of a field: "Always block-card when topic is lost_stolen and card_status is active or frozen", "Never answer-faq when topic is unrecognised". Say them to the set-up agent like any other rule, or write the condition as `{ "answer": "topic", "op": "eq", "value": "lost_stolen" }`. What a message is about is never certain, so these rules are **enforced whenever there's a real chance they apply**, not only when it's the likeliest reading: - The rule applies when the chance it holds is at least its `min_chance` (0.15 unless you set another, from 0 to 1). - When it applies but isn't sure to (a chance under 0.5) and it changes the action, the decision goes to your **review** queue instead of being acted on. - The decision lists each such rule's chance in `rule_chances`, and the decision log shows it next to the rule. - It works the same hosted and in a downloaded model, and it reads the answer it needs even when you only ask for the decision. A rule on a field is exact; if your system records the fact ("card_reported_lost is true"), a rule on that field is the strongest form. ## Act, review, ask Each model sets its own **safety bar** after every training run: the lowest confidence at which its automatic answers are at least 97% right on your held-back cases. There's nothing to set. The bar turns confidence into a route: | When | `route` | What to do | |---|---|---| | at or above the bar, in a language the model is strong in | `act` | act on it | | below the bar (or another language), still below after it was double-checked | `review` | hand it to a person; it joins your [unsure queue](https://canonopylabs.com/docs/improving.md#the-unsure-queue) | | your rules allow no action | `ask` | hand it to a person; it's in the unsure queue too | A double-checked message is translated to English and decided again; if that's still unsure, your day-one backend answers (Jev or an OpenAI-compatible model) or, without one, our hosted backup; if neither can answer, it's held for a person. See [Languages](https://canonopylabs.com/docs/languages.md) and [Confidence](https://canonopylabs.com/docs/confidence.md). ## The decision log Every decision is kept, so you can look back at what happened: to debug an integration, or for your own records. In the console, open a decision model and choose **Decisions**; or use the API, the CLI, the SDK or your coding agent. - **The table:** newest first, with the time, a short summary of the case, the final decision, its confidence, where it came from (`source`), the `route` and the version. - **Filters:** a time range, `source` (`trained`, `translated`, `fallback`, `rule`, `backup`, `local_fallback`), `route`, an answer (`action:refund`), the version, and whether an outcome was reported. Search looks for words in the case's text, or a decision id. - **One decision in full:** the case as you sent it, every answer with its probabilities and confidence, the final decision and the rules of yours that blocked options ("Blocked by your rule: never approve over 500"), the version, the detected language, how long it took, and the outcome once reported. Cases your [downloaded models](https://canonopylabs.com/docs/running-it-yourself.md#save-and-sync-later) decided offline and synced later are marked. - **Mark wrong:** give the right answer for a decision (**Mark wrong** in the console, or `POST /v1/decisions/{decision_id}/wrong`). It's recorded as its outcome, marked in the log, and the next retrain learns it. See [Mark a decision wrong](https://canonopylabs.com/docs/improve-your-model.md#mark-a-decision-wrong). - **Export:** CSV or JSONL with the same filters, up to 100,000 decisions per file. - **Deleting:** one decision, or every decision before a date (confirmed with the model's name). Deleting is permanent. The log stays readable after your free trial ends: it's your data, like your downloads. ```bash curl "https://api.canonopylabs.com/v1/domains/refunds/decisions?route=review&since=2026-09-01&answer=action:approve" \ -H "Authorization: Bearer $CANONOPY_API_KEY" curl "https://api.canonopylabs.com/v1/domains/refunds/decisions/export?format=csv&since=2026-09-01" \ -H "Authorization: Bearer $CANONOPY_API_KEY" -o refunds-decisions.csv ``` ### How long decisions are kept Decisions are kept for **90 days** unless you change it in the console's Settings (or `PATCH /v1/settings`): 30, 90 or 365 days, or until you delete them. Once a day, older decisions are deleted permanently in every decision model of the workspace. That includes decisions with an outcome or an answer from the unsure queue, and open cases in the queue. Once they're deleted, later versions no longer learn from them, and a question still in [day-one mode](https://canonopylabs.com/docs/day-one.md#graduation) loses the outcomes it had counted. Versions already trained keep what they learned. If you rely on your outcomes to keep improving, keep decisions for 365 days or until you delete them. --- Source: https://canonopylabs.com/docs/confidence (Markdown: https://canonopylabs.com/docs/confidence.md) # Confidence Every Choice and Score answer has a `confidence`; a Noul is a probability itself. They're **calibrated**: across many answers where it says 90%, it's right about 90% of the time. ## Measured on your cases, not promised Calibration is measured across groups of answers, not guaranteed for any single one. Every training report shows yours, on your own held-back cases: - a **calibration error** (0 is perfect; the bank test scored 0.006, against 0.044–0.063 for Jev); - a **reliability curve**: stated confidence against how often it was right; - a sentence you can quote, like "When it says 90–100% sure, it's right 99.3% of the time." ## The safety bar sets itself You don't choose a cut-off. After every training run, the model sets its own **safety bar**: the lowest confidence at which its automatic answers are at least **97% right** on your held-back cases, checked language by language. Each retrain measures it again, and the report says it in plain words: "automatic answers 97%+ accurate; 78% handled automatically; the rest double-checked". *Interactive on the web page (https://canonopylabs.com/docs/confidence): a chart of confidence against accuracy, and of the share handled automatically at each safety bar.* Below the bar, or in a language the model isn't strong in, a message is double-checked: translated to English and decided again, then your day-one backend, else held for a person. See [Languages](https://canonopylabs.com/docs/languages.md). ## Using confidence in code ```python d = res["decision"] if d["route"] == "act": run(d["action"]) # automatic: at or above the safety bar else: hand_to_person(res["id"]) # review or ask: it's also in your unsure queue ``` - **Prefer `route`** over raw numbers: it follows the model's safety bar, which moves with every retrain. Code written for Jev that reads `confidence` keeps working. - **Check `source`.** A `fallback` answer's confidence comes from your day-one backend, calibrated its own way; a `backup` answer's comes from our hosted backup. - **Middling scores need a look at `confidence`.** A Score of 1.5 can mean "clearly in the middle" or "torn between the ends". See the explorer on [Questions](https://canonopylabs.com/docs/questions.md#score). --- Source: https://canonopylabs.com/docs/languages (Markdown: https://canonopylabs.com/docs/languages.md) # Languages Every model **reads about 100 languages**; accuracy is measured in 51. A customer can write in Spanish, Hindi or Japanese, and the same model decides, with your rules enforced. ## Strong languages Each version knows which languages it's strong in. Its training report lists them, with results per language, and so does the domain (`safety_bar.strong_languages`). A strong language is one where the model's automatic answers clear the safety bar on your held-back cases. ## The safety bar: 97%, set for you Automatic answers are held to a fixed bar: **at least 97% right**. After every training run, the model sets the confidence each answer needs to reach it, from your own held-back cases, and it's checked language by language. Each retrain measures it again. There's no threshold to tune. ## Unsure messages are double-checked Two checks run in code on every answer, never left to the model: - **Language:** a message in a language the model isn't strong in is always double-checked. - **Confidence:** an answer below the bar is double-checked. A double-checked message is translated to English and the same model decides again, with your rules still applied. That answer says `"source": "translated"`. If it's still below the bar, your day-one backend (Jev or an LLM) answers (`"source": "fallback"`). Without one, our hosted backup answers from your allowed options, with your rules still applied (`"source": "backup"`). If the backup is unavailable or can't tell, it's held for a person: `"route": "review"`, and it waits in your [unsure queue](https://canonopylabs.com/docs/improving.md). ```json { "answers": { "action": { "type": "choice", "choice": "block-card", "confidence": 0.98, "probabilities": { "block-card": 0.98, "open-dispute": 0.02 }, "source": "translated", "route": "act" } }, "decision": { "question": "action", "action": "block-card", "confidence": 0.98, "blocked_by_rules": [], "blocked_actions": [], "route": "act" }, "language": "es" } ``` `language` is the language the message was written in. It's `null` when there's too little text to tell (a number, or "ok"): the confidence check still applies. ## A weak language gets better with cases When a language is below the bar, the training report says so and suggests what fixes it, for example "Upload ~300 cases in Swahili to handle more automatically". Past cases written in that language, with how they were resolved, raise the share it handles on its own. ## Running it yourself A [downloaded model](https://canonopylabs.com/docs/running-it-yourself.md) runs the same two checks. With your API key set, it hands a double-checked message to us. Offline, it holds the message for review (`"route": "review"`). It never guesses. --- Source: https://canonopylabs.com/docs/day-one (Markdown: https://canonopylabs.com/docs/day-one.md) # Day one: answers from your first call You don't need history to start. The quickest way to your own model is to [send the JSON you send Jev](https://canonopylabs.com/docs/start-with-no-data.md): it's ready in minutes. Day-one mode covers the time before that, and any question you haven't built a model for: a new domain, or a question it hasn't learned yet, is answered by **your day-one backend**, in the same shape, marked `"source": "fallback"`. Your rules still apply through the `decision` block, and so does the safety check: a day-one answer below 0.8 confidence comes back `"route": "review"` and waits in the [unsure queue](https://canonopylabs.com/docs/improving.md#the-unsure-queue) instead of being acted on. When a question has enough of **your real outcomes** (from uploads, `/v1/outcomes` and the unsure queue), it **graduates to your own model** automatically, question by question. Each graduation makes that question faster, cheaper to run and, in our tests, more accurate. ## A starter model for decisions on text A model you [build from your Jev request](https://canonopylabs.com/docs/start-with-no-data.md) reads text from day one, with no past cases. If you set up by describing the decision to the set-up agent instead, a question that reads text, such as what a message is about, or a routing decision like *billing, shipping or account* that depends on what the message says, can **start with a starter model**: train, and the question has a working model of your own on day one, with your rules enforced on every decision. There's nothing to set or ask for; it's used when it helps. - **Your rules still apply on top.** A rule that says never still blocks, and a rule on your fields (like "escalate when amount is over 1,000") still decides the cases it covers. - **Say what each option covers** when you describe the decision, like `Question topic: billing (payments, invoices), shipping (delivery, tracking), account (login, password)`: it helps the starter model and day-one answers tell them apart. - **It learns your wording** as you [upload past cases](https://canonopylabs.com/docs/training-and-reports.md#upload-history) or [report outcomes](https://canonopylabs.com/docs/improving.md#report-outcomes). Each category moves over to your own cases once it has enough of them, and your recorded answers always win. - **Its accuracy is labelled.** Until you have enough cases of your own, the report measures it on starter example cases (`measured_on: "starter_cases"`) and says so: *"Starter model: ready to use now. It learns your wording as you upload past cases or report outcomes; real accuracy appears once there are enough of them."* From then on every number comes from your own held-back cases, and the [never-worse check](https://canonopylabs.com/docs/training-and-reports.md) compares versions on those. - **Its automatic answers are more cautious** while it's measured on starter cases: more of them are double-checked. - **You can remove the starter cases** any time with `DELETE /v1/domains/{domain}/uploads/starter` (or **Remove the starter cases** on the Data tab). They then stay off for that model. Decisions on data (numbers and fields only) don't need one: for them there's [nothing to bring](https://canonopylabs.com/docs/what-you-need-to-bring.md). ## Coming from Jev? The request and answer shapes are the same, so switching is a change of URL and `model`. ```diff - POST https://api.typesafe.ai/v1/systemone "model": "jev-latest" + POST https://api.canonopylabs.com/v1/systemone "model": "support@latest" ``` A domain that doesn't exist yet is created in day-one mode on its first call, with the questions that call asks. A `model` that starts with `jev-` maps to your `default` domain. ## Questions your model doesn't have Ask a question your model doesn't have and it is answered all the same (`"source": "fallback"`). It doesn't join your model on its own: it's listed in the model's `candidate_questions`, and requests without `questions` keep asking only your model's own questions. Outcomes you report for it are kept. When you want it for good, add it to the model's `questions` with `PATCH /v1/domains/{domain}`; from then on it counts toward graduation like any other question. ## Choose your backend **`PUT /v1/domains/{domain}/fallback`** | `provider` | What answers | |---|---| | `typesafe` | Jev, with your TypeSafe key | | `openai` | any OpenAI-compatible endpoint (`base_url`, `model`, key) | | `local` | our built-in general model: no key, no cost | | `none` | no day-one backend | ```json { "provider": "typesafe", "api_key": "", "model": "jev-latest" } ``` Keys are stored encrypted, used only for that domain's fallback calls, and never returned: reads show `has_key` and the last four characters. ## Offline: your own fallback, and sync later A [downloaded model](https://canonopylabs.com/docs/running-it-yourself.md) can't call your day-one backend while it's offline. Give it a function of your own instead (your rules of thumb, or a small model you run): `load("refunds@3.zip", offline_fallback=my_fallback)`. It answers the questions that need a second look when we can't be reached, marked `"source": "local_fallback"`, with your rules still applied after it; `None` means review. The runtime also saves those cases on your machine, and `model.sync()` (or `canonopy sync`) sends them to your unsure queue later, so they count toward graduation like any other outcome once answered. See [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md#your-own-fallback). ## Graduation **`GET /v1/domains/{domain}/graduation`** ```json { "min_outcomes": 1000, "auto": true, "questions": { "action": { "status": "trained", "version": 3, "outcomes": 5120, "needed": 0 }, "urgency": { "status": "fallback", "version": null, "outcomes": 213, "needed": 787 } } } ``` Outcomes come from your [uploads](https://canonopylabs.com/docs/training-and-reports.md#upload-history), from [`POST /v1/outcomes`](https://canonopylabs.com/docs/improving.md#report-outcomes) and from answers in the [unsure queue](https://canonopylabs.com/docs/improving.md#the-unsure-queue). Set `graduation: {"min_outcomes": 1000, "auto": true}` on the domain; with `auto: false`, you decide when to train. Graduation is checked each time an outcome or an unsure-queue answer comes in, so after an upload that takes a question past the mark, it graduates with the next one; or train it straight away with `POST /v1/domains/{domain}/train`. > **What day one can't do** > Until a question graduates, its accuracy is your backend's accuracy. The console shows how many outcomes each question still needs. --- Source: https://canonopylabs.com/docs/with-your-history (Markdown: https://canonopylabs.com/docs/with-your-history.md) # Make it better with your history No data: your model is ready in minutes. With your decisions: the same builder makes it better. On 3,080 real bank messages, models built with no data scored 84.6–86.8% against Jev's 86.6%; after about 400 reviewed cases, 91.7%; with the bank's history, about 95% (see [Bank support, from day one](https://canonopylabs.com/docs/cookbooks/bank-zero-data.md)). Past cases with the answers you gave them are used automatically, in any build of a model made with [`POST /v1/build`](https://canonopylabs.com/docs/start-with-no-data.md): its first build, a rebuild under the same name, or a retrain. ## Send your past cases Two ways, same result: - **In the request:** add `history` to `POST /v1/build`, up to 50,000 cases. - **As an upload:** upload them to the model (`POST /v1/domains/{domain}/data`, a CSV, JSONL or JSON file, see [Upload history](https://canonopylabs.com/docs/training-and-reports.md#upload-history)), then build again under the same name or [retrain](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain). Each case is `{"state", "answers"}`, the state exactly as your program sends it, or a flat row of fields and answers, exactly as in an upload. **curl** ```bash curl https://api.canonopylabs.com/v1/build \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "state": { "order": { "amount": 42.5 }, "customer": { "tier": "pro", "fraud_flag": false } }, "questions": { "action": { "type": "choice", "instructions": "What should we do with this refund request?", "criteria": { "approve": "refund the customer", "escalate": "a person decides", "decline": "politely decline" } } }, "name": "refunds", "history": [ { "state": { "order": { "amount": 820 }, "customer": { "tier": "pro", "fraud_flag": false } }, "answers": { "action": "escalate" } }, { "state": { "order": { "amount": 35 }, "customer": { "tier": "free", "fraud_flag": false } }, "answers": { "action": "approve" } } ] }' ``` **Python** ```python import json from canonopy import Client client = Client() request = json.load(open("jev-request.json")) # the body you send Jev past = [json.loads(line) for line in open("past_refunds.jsonl")] # {"state": ..., "answers": ...} per line build = client.build(state=request["state"], questions=request["questions"], name="refunds", history=past) build = client.wait_for_build(build["build_id"]) print(build["history"]["statement"]) ``` **CLI** ```bash canonopy build jev-request.json --name refunds --history past_refunds.csv # CSV, JSONL or JSON # or, for a model you already built: canonopy upload refunds past_refunds.csv canonopy train refunds # a retrain uses them ``` `understood.history` says how many were stored (`accepted`, `rejected`, and the `problems` with the first few). They're stored as one upload of the model: listed, replaceable and deletable like any upload. ## What your history changes - **Your cases are the main examples.** Your model learns from your real cases and their answers first. For text fields, it also learns from variations of your messages: typos, slang, other wording. - **Situations fill the gaps.** Answers your history barely has, and rare cases it doesn't cover, are filled with situations for your questions. `history.gaps` lists them: `{"question": "action", "answer": "decline", "your_cases": 12, "situations": 900}`. - **Rules found in your history are proposed, not used.** Each comes with how many past cases it applies to and how often they agree with it. You confirm or reject each one; only confirmed rules are used. See [Rules found in your history](https://canonopylabs.com/docs/rules-found.md). - **Cases that contradict your rules are flagged.** A past case that goes against one of your hard rules, or a rule found in your history that you confirmed, is never learned silently: it isn't learned at all until you review it. See [Flagged past cases](https://canonopylabs.com/docs/rules-found.md#flagged-past-cases). - **It's measured on your own cases.** About 15% of your past cases (at most 2,000) are held back, never learned from, and your model is measured on them. That's your honest real-world number. ## Measured on your own held-back cases After the usual [quality check on fresh situations](https://canonopylabs.com/docs/start-with-no-data.md#the-quality-check-95-on-fresh-situations), your model is compared with the answers you gave your held-back cases: ```json "own_cases": { "cases": 600, "questions": [{ "question": "action", "agreement": 0.962, "cases": 600 }], "statement": "Measured on your own cases: your model agrees with your past answers in 96.2% or more of 600 held-back cases it never learned from, for every question." } ``` - It's in the build's `quality.own_cases`, and the version's report shows the same numbers with `measured_on: "your_cases"`. - **The same cases are held back on every retrain** (plus a share of your new ones), so versions are compared on the same cases. - With fewer than 40 past cases nothing is held back, and `own_cases` is `null`. - Serving still needs the usual 95% bar on fresh situations. Where your past answers differ from your instructions and rules, the number on your own cases shows it. ## How your cases were used The build's `history`: ```json "history": { "cases": 4200, "learned_from": 3558, "held_back": 630, "variations": 2400, "flagged": 12, "pending": 12, "gaps": [{ "question": "action", "answer": "decline", "your_cases": 12, "situations": 900 }], "statement": "…" } ``` | field | meaning | |---|---| | `cases` | your past cases with an answer | | `learned_from` | the ones your model learns from: not held back, and not flagged and waiting for you | | `held_back` | held back to measure your model on your own cases | | `variations` | variations of your messages it also learns from | | `flagged`, `pending` | cases that contradict a rule, and those still waiting for your review | | `gaps` | answers your history rarely has, filled with situations | ## The flow 1. Build with `history` (or upload, then build again or retrain). Your model trains, is checked on fresh situations and measured on your own cases, with nothing to do in between. 2. Optional, once it's ready: [confirm or reject the rules found](https://canonopylabs.com/docs/rules-found.md), [review the flagged cases](https://canonopylabs.com/docs/rules-found.md#flagged-past-cases) (until then they aren't learned), and review how it decides on about 20 examples. Your decisions and corrections apply at the next retrain. 3. Later, upload more cases and [retrain](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain): the same held-back cases, so you see real progress. In the console, the build page shows the rules found and the flagged cases, and the model page's **Review examples** shows how it decides. With a coding agent, the MCP tools are `build_model` (with `history` rows, or `history_file` on the local server), `get_build`, `confirm_found_rules` and `review_flagged_history`. In the CLI: `canonopy build … --history FILE`, `canonopy rules found` and `canonopy history flagged`. --- Source: https://canonopylabs.com/docs/rules-found (Markdown: https://canonopylabs.com/docs/rules-found.md) # Rules found in your history When you build with [your past cases](https://canonopylabs.com/docs/with-your-history.md), we look for the rules your history already follows and propose them to you, each checked on every past case. **Only the ones you confirm are used.** Past cases that contradict your rules are flagged for you to review, and aren't learned until you do. ## What a found rule shows They're in the build's `found_rules`, strongest first, at most 10: ```json { "id": "fr_1", "sentence": "Escalate refunds over 500", "question": "action", "answer": "escalate", "support": 4000, "agreement": 0.97, "status": "proposed", "hard": false, "can_be_hard": true, "statement": "Escalate refunds over 500: true in 97% of the 4,000 past cases it applies to.", "changes": { "situations": 120, "of": 30000, "statement": "Changes 120 of 30,000 situations: approve → escalate." } } ``` | field | meaning | |---|---| | `sentence` | the rule in plain words | | `question`, `answer` | the question it answers, and the answer it gives where it applies | | `support` | your past cases it applies to | | `agreement` | how often those cases had its answer | | `changes` | what confirming it changes in the situations your model learns from | | `status` | `proposed` (waiting for you: not used), `confirmed` or `rejected` | | `hard`, `can_be_hard` | confirmed as a hard rule; whether it can be one | Only rules your history really shows are proposed: at least 5 cases and 0.5% of them, with at least 80% agreement. ## Check each one A rule found in your history is what you **did**, not necessarily what you **want**. Past habits get copied too: a workaround, a policy you've since changed, one team's shortcut. So nothing is used until you decide: - **Confirm:** your model follows it. - **Confirm as hard** (when `can_be_hard`: it answers the Choice question your hard rules apply to): it's also enforced on every decision, in code, like your own rules, from the retrain it starts. - **Reject:** it isn't used. `changes` tells you what confirming it does before you do it. A rule with 97% agreement also means 3% of those cases went the other way: if they were right, reject it, or confirm it and keep those cases when they're [flagged](#flagged-past-cases). **`POST /v1/build/{build_id}/found-rules`** **curl** ```bash curl https://api.canonopylabs.com/v1/build/bld_…/found-rules \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"rules": [{"id": "fr_1", "decision": "confirm", "hard": true}, {"id": "fr_2", "decision": "reject"}, {"id": "fr_3", "decision": "confirm"}]}' ``` **Python** ```python build = client.latest_build("refunds") for r in build["found_rules"]: print(r["id"], r["statement"], r["changes"]["statement"]) client.confirm_found_rules(build["build_id"], [ {"id": "fr_1", "decision": "confirm", "hard": True}, {"id": "fr_2", "decision": "reject"}, {"id": "fr_3", "decision": "confirm"}, ]) ``` **CLI** ```bash canonopy rules found refunds # list them canonopy rules found refunds --confirm fr_3 --hard fr_1 --reject fr_2 # --hard confirms and enforces it ``` It answers with the `Build`. This is optional, and nothing waits for it: your model is trained and serves without the found rules, and your decisions apply from the next [retrain](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain), with a note that says so. `400`: an unknown id, or `hard` on a rule that can't be hard. `409`: the build is busy (queued, building, training or checking); try again when it has finished. ## Flagged past cases A past case that contradicts one of your hard rules, or a found rule you confirmed, is flagged. **It's never learned silently: until you review it, it isn't learned at all.** The build's `flagged` has the counts and the first 20: ```json "flagged": { "total": 12, "pending": 12, "items": [ { "id": "h1042", "state": { "order": { "amount": 820 }, "customer": { "tier": "pro", "fraud_flag": false } }, "answers": { "action": "approve" }, "rule": "Escalate refunds over 500", "question": "action", "rule_answer": "escalate", "kind": "found", "status": "pending", "why": "…" } ] } ``` `kind` is `found` (a rule found in your history you confirmed) or `hard` (one of your hard rules). `rule_answer` is the rule's answer, or `null` for a rule that only blocks answers. **`GET /v1/build/{build_id}/flagged`** All of them, a page at a time: `status` (`pending` or `reviewed`; default all), `limit` (1-200, default 50) and `cursor` (the `next_cursor` of the page before). → `build_id`, `domain`, `total`, `pending`, `items`, `next_cursor`. **`POST /v1/build/{build_id}/flagged`** For each case: | decision | what happens | |---|---| | `follow_rule` | learn it with the rule's answer (not for a rule that only blocks answers) | | `keep` | learn it as it is: the rule has an exception | | `drop` | never learn it | **curl** ```bash curl https://api.canonopylabs.com/v1/build/bld_…/flagged \ -H "Authorization: Bearer $CANONOPY_API_KEY" -H "Content-Type: application/json" \ -d '{"rows": [{"id": "h1042", "decision": "follow_rule"}, {"id": "h2210", "decision": "drop"}]}' # or one decision for every case still pending: curl https://api.canonopylabs.com/v1/build/bld_…/flagged \ -H "Authorization: Bearer $CANONOPY_API_KEY" -H "Content-Type: application/json" \ -d '{"all": "follow_rule"}' ``` **Python** ```python page = client.flagged_history(build_id, status="pending") for row in page["items"]: print(row["id"], row["answers"], "vs", row["rule"], row["why"]) client.review_flagged_history(build_id, rows=[{"id": "h1042", "decision": "follow_rule"}, {"id": "h2210", "decision": "drop"}]) client.review_flagged_history(build_id, all="follow_rule") ``` **CLI** ```bash canonopy history flagged refunds --status pending # list them canonopy history flagged refunds --follow-rule h1042 --keep h1180 --drop h2210 canonopy history flagged refunds --all follow_rule ``` → `reviewed`, `total`, `pending` and a plain-words `message`. With `all: "follow_rule"`, cases whose rule only blocks answers stay pending: decide those one by one. **Your decisions apply when the model is next trained**, at the next [retrain](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain). Reviewing them is optional: nothing waits on it. ## Everywhere | | found rules | flagged cases | |---|---|---| | API | `POST /v1/build/{build_id}/found-rules` | `GET` and `POST /v1/build/{build_id}/flagged` | | CLI | `canonopy rules found [--confirm …] [--reject …] [--hard …]` | `canonopy history flagged [--status …] [--follow-rule …] [--keep …] [--drop …] [--all …]` | | MCP | `confirm_found_rules` (`build_id` or `model`, `rules`: each `id`, `decision`, `hard`) | `review_flagged_history` (`build_id` or `model`; with no decisions it lists the pending ones; `rows` or `all`) | | Python | `client.confirm_found_rules(build_id, rules)` | `client.flagged_history(build_id, …)`, `client.review_flagged_history(build_id, rows=…, all=…)` | | Console | the build page: **Rules found**, confirm, reject or hard, each with what it changes | the build page: **Flagged history** | A coding agent should show each rule and its `changes` to you and wait for your decision: confirming a rule changes what your model learns, and a hard rule changes what it's allowed to answer. --- Source: https://canonopylabs.com/docs/improve-your-model (Markdown: https://canonopylabs.com/docs/improve-your-model.md) # Improve your model Two things can be off. Your model may not match your instructions and rules somewhere: a **targeted retrain** sharpens it there. Or a rule itself may be wrong: **change the rule**: it retrains straight away, and the build says what it changed. A retrain brings your model closer to your instructions and rules; it can't make a wrong rule right, so the advice below points out rules your outcomes contradict. **Every retrain is your call.** Nothing retrains, promotes or changes a rule on its own. This page is about models made with [`POST /v1/build`](https://canonopylabs.com/docs/start-with-no-data.md). Advice and marking a decision wrong work for every model. For outcomes, the unsure queue, new options and replacing cases, see [Improving](https://canonopylabs.com/docs/improving.md). ## Advice **`GET /v1/domains/{domain}/advice`** What would improve your model most right now, in plain words, most useful first. It's also in `GET /v1/domains/{domain}/progress` as `advice`, and on the console's **Progress** tab as cards. Read-only. ```json { "domain": "cards", "statement": "3 things would improve your model most right now.", "items": [ { "kind": "retrain_to_sharpen", "title": "Retrain to sharpen 'topic'", "action": "retrain", "question": "topic", "agreement": 0.91, "text": "'topic' matches your instructions and rules in 91% of situations when the right answer is 'lost_card' (97.8% overall). A retrain adds situations where your model disagrees and shows the before and after for them." }, { "kind": "rule_may_be_wrong", "title": "This rule may be wrong", "action": "change_rules", "rule": "Never approve when amount is over 500", "contradicted": 7, "applied": 20, "text": "…in 7 of the 20 outcomes where it applied, the right answer went against it ('approve'). If it should allow that, change the rule: the build shows what it changes." }, { "kind": "answer_these", "title": "Answer a few of these", "action": "answer", "decisions": [{ "decision_id": "dec_…", "summary": "…", "answer": "refund", "confidence": 0.41 }], "text": "These 5 decisions in your unsure queue are the ones your model was least sure about. …" } ] } ``` | `kind` | when | `action` | |---|---|---| | `retrain_to_sharpen` | a built model's latest check is under 99% somewhere | `retrain`: a [targeted retrain](#targeted-retrain) | | `rule_may_be_wrong` | at least 3 reported outcomes, unsure-queue answers or decisions marked wrong, and 20% of those where the rule applied, went against one of your hard rules or a confirmed [rule found in your history](https://canonopylabs.com/docs/rules-found.md) | `change_rules`: [change it with a preview](#change-a-rule-with-a-preview) | | `answer_these` | decisions whose answers would teach the next retrain most: your unsure queue, else the least sure recent decisions | `answer`: the [unsure queue](https://canonopylabs.com/docs/improving.md#the-unsure-queue), or [mark them wrong](#mark-a-decision-wrong) | | `upload_these` | a built model with no past cases of its own | `upload`: see [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md) | CLI: `canonopy advice refunds`. MCP: `get_advice` (`model`). Python: `client.advice("refunds")`. ## Targeted retrain **`POST /v1/domains/{domain}/train`** The same call as any retrain. On a model made with `POST /v1/build`, it runs as a build of kind `retrain`, from the same instructions and rules, and adds: - situations where the current version disagrees with your instructions and rules; - more around the answers it confuses, and around the answers your hard rules depend on and their look-alikes (for "block the card when the topic is lost_card": a compromised card, a stolen phone); - your new cases since the last version (uploads, outcomes, unsure-queue answers); - the decisions you [marked wrong](#mark-a-decision-wrong), and the examples you [corrected](https://canonopylabs.com/docs/start-with-no-data.md#review-how-it-decides-optional). Your instructions are made to give those right answers first. Nothing waits for you. **curl** ```bash curl -X POST https://api.canonopylabs.com/v1/domains/cards/train \ -H "Authorization: Bearer $CANONOPY_API_KEY" -H "Content-Type: application/json" -d '{"wait": false}' # → a Job with build_id; follow it: curl https://api.canonopylabs.com/v1/build/bld_… -H "Authorization: Bearer $CANONOPY_API_KEY" ``` **Python** ```python job = client.train("cards", wait=False) build = client.wait_for_build(job["build_id"]) print(build["retrain"]["statement"]) ``` **CLI** ```bash canonopy train cards # prints the build id canonopy build status cards # where it is, then the before and after ``` The `Job` has `build_id`; with `wait: true` it waits for the build and answers with its training job (the version and its report). The build's `retrain` says what it added and, once checked, the before and after on the same fresh situations: ```json "retrain": { "since_version": 3, "added": { "disagreements": 312, "around_confusions": 900, "look_alikes": 600, "example_messages": 40, "corrections": 2 }, "areas": [{ "question": "topic", "answer": "lost_card", "area": "when the right answer to 'topic' is 'lost_card'", "before": 0.91, "after": 0.99, "cases": 240 }], "statement": "Added 312 situations where version 3 disagreed with your instructions and rules and 1,500 more around 'lost_card', 'card_fault' and their look-alikes. Before and after: when the right answer to 'topic' is 'lost_card': 91% → 99%." } ``` The version's report has it too (`build.retrain`), and the **Progress** tab compares the versions. The new version serves when every question passes the [95% bar](https://canonopylabs.com/docs/start-with-no-data.md#the-quality-check-95-on-fresh-situations); with your history, it's also measured on the [same held-back cases of yours](https://canonopylabs.com/docs/with-your-history.md#measured-on-your-own-held-back-cases) as before. `questions` in the body retrains only those questions, the usual way. A retrain counts as a build (20 per workspace per UTC day). ## Change a rule with a preview **`POST /v1/domains/{domain}/rules`** Say the change in plain words. Your model is retrained with the new rules straight away, and the new version serves only if it passes the [quality check](https://canonopylabs.com/docs/start-with-no-data.md#the-quality-check-95-on-fresh-situations); until then, `@latest` keeps the version before. ```json { "rules": "Always escalate when amount is over 300.", "mode": "add" } ``` `mode`: `add` (to your current rules, the default) or `replace` (these are all your rules now). It answers with a `Build` of kind `rule_change`. Its `rule_preview` says what changes: ```json "rule_preview": { "rules": "Never approve when amount is over 500.\nAlways escalate when amount is over 300.", "previous_rules": "Never approve when amount is over 500.", "situations": 1240, "of": 3000, "changes": [{ "question": "action", "from": "approve", "to": "escalate", "situations": 1100 }, "…"], "examples": [{ "state": { "…": "…" }, "before": { "action": "approve" }, "after": { "action": "escalate" } }, "…"], "statement": "Changes 1,240 of 3,000 situations: approve → escalate (1,100), decline → escalate (140)." } ``` About 10 changed situations come with it. The build's examples (`GET /v1/domains/{domain}/examples`) are situations whose answer changes: reviewing them is optional. If one is wrong, correct it with `POST /v1/domains/{domain}/signoff` and `"retrain": true`, and the next version gives that answer. **curl** ```bash curl https://api.canonopylabs.com/v1/domains/refunds/rules \ -H "Authorization: Bearer $CANONOPY_API_KEY" -H "Content-Type: application/json" \ -d '{"rules": "Always escalate when amount is over 300."}' ``` **Python** ```python build = client.change_rules("refunds", "Always escalate when amount is over 300.") # mode="add" build = client.wait_for_build(build["build_id"]) print(build["rule_preview"]["statement"]) ``` **CLI** ```bash canonopy rules change refunds "Always escalate when amount is over 300." # --replace for all your rules canonopy build status refunds # what changed, then the quality check canonopy examples refunds # optional: the changed situations ``` A message to the set-up agent (`POST /v1/domains/{domain}/agent`) on a built model does the same with `mode: add`; its reply has the `build_id`. MCP: `change_rules` (`model`, `rules`, `mode`). In the console, change the rules on the model: it retrains, and the build shows what changed. Errors: `400` (not a model made with `POST /v1/build`: change its rules with the set-up agent; or no rules), `402`, `409` (a build is in progress), `429` (a rule change counts as a build). ## Mark a decision wrong **`POST /v1/domains/{domain}/decisions/{decision_id}/wrong`** **`POST /v1/decisions/{decision_id}/wrong`** The second works with the decision id alone, as `POST /v1/decide` returned it. ```json { "answers": { "action": "escalate" } } { "action": "escalate", "note": "over the limit" } ``` `answers` (per question) or `action` (the decision question's answer), and an optional `note` for your own records. → `{"recorded": true, "decision_id", "domain", "answers", "message"}`. - It's recorded as the decision's outcome, with `via: "marked_wrong"`; the decision log shows it (`marked_wrong: true`), and **Progress** counts it under "Decisions you marked wrong". - **The next retrain learns it.** On a built model, the retrain first makes your instructions give that answer there. - When your corrections go against a rule, [advice](#advice) says so (`rule_may_be_wrong`). CLI: `canonopy decisions wrong dec_… --answer escalate` (or `--answer action=escalate`, and `--note "…"`). MCP: `mark_decision_wrong` (`decision_id`, `answers` or `action`, `note`). Python: `client.mark_wrong("dec_…", action="escalate")`. Console: **Mark wrong** on a decision in the [decision log](https://canonopylabs.com/docs/decisions-and-rules.md#the-decision-log). `400`: an unknown question or answer; `404`: no such decision. ## A loop that works 1. Read the advice on **Progress** (or `canonopy advice`). 2. Weak spot? Retrain, and read the before and after. 3. A rule your outcomes contradict? Change it, and read what it changed (the changed examples are optional to review). 4. Mark the wrong decisions you find in the log, and answer the unsure queue: the next retrain learns them. 5. Promote or roll back as usual: see [Versions](https://canonopylabs.com/docs/training-and-reports.md#versions). --- Source: https://canonopylabs.com/docs/playtest-your-model (Markdown: https://canonopylabs.com/docs/playtest-your-model.md) # Playtest your model Before your model plays for real, let your own game play with it. `canonopy playtest` runs on your machine: the only thing of ours that runs there is the finished model you downloaded, and only the numbers your program prints come back. ```bash canonopy playtest my-bot@latest --cmd "python my_game.py" --episodes 20 ``` It needs `canonopy-runtime` (see [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md)). ## How your game talks to the model 1. The model is served at a small local endpoint, on `127.0.0.1` only. Your program sends the same JSON you send Jev and gets the usual answer: ```python import json, os, urllib.request def decide(state): req = urllib.request.Request(os.environ["CANONOPY_DECIDE_URL"], data=json.dumps({"state": state}).encode(), headers={"Content-Type": "application/json"}) return json.loads(urllib.request.urlopen(req).read())["answers"] ``` 2. Your command runs once per episode (or once for all of them with `--once`), with `CANONOPY_DECIDE_URL`, `CANONOPY_EPISODE` (1, 2, …), `CANONOPY_EPISODES` and `CANONOPY_SEED` set. 3. At the end of each episode it prints **one JSON line** of numbers: anything you measure. ```python print("CANONOPY_RESULT " + json.dumps({"score": 1240, "deaths": 1, "stuck": 0})) ``` The `CANONOPY_RESULT ` prefix is optional: without it, the episode's last line that is a JSON object counts. Everything else your program prints is shown, untouched. ## What comes back The numbers go to your workspace (`POST /v1/domains/{domain}/play-results`), summarised per number: the average, the lowest and the highest. The advice shows your last playtest, and you can compare after a retrain. **`POST /v1/domains/{domain}/play-results`** **`GET /v1/domains/{domain}/play-results`** **Improve from play (optional):** `--keep-situations 200` also sends a sample of up to that many states your game sent to the model (300 at most). Your next retrain answers them with your model's own instructions and learns them, so it gets better where your game really goes. | flag | what it does | |---|---| | `--episodes N` | episodes to play (default 10) | | `--once` | run your command once; it plays every episode and prints a line for each | | `--timeout S` | seconds per episode (default 600) | | `--port P` | the local port (default: any free one) | | `--keep-situations N` | also send a sample of the states your game hit | | `--no-send` | print the results, send nothing | | `--name MODEL` | with a local `.zip`: the model to send the results to | Without the CLI: `canonopy-runtime playtest MODEL.zip --cmd "..." --episodes 20` prints the same results, and your coding agent can send them with `report_play_results`. --- Source: https://canonopylabs.com/docs/shadow-mode (Markdown: https://canonopylabs.com/docs/shadow-mode.md) # Shadow mode Try your model on real work before it acts. In shadow mode it runs beside Jev (or your current process) on real requests without acting: you keep deciding as today, it decides silently, and you see where the two disagree. ```bash canonopy shadow start refunds canonopy shadow send refunds requests.jsonl # {"state": …, "current": {"action": "approve"}, "source": "jev"} per line canonopy shadow review refunds # a short list of disagreements ``` **`PUT /v1/domains/{domain}/shadow`** **`POST /v1/domains/{domain}/shadow/cases`** Send up to 100 requests per call, each with the decision you made today: ```json { "cases": [ { "state": { "order": { "amount": 42.5 }, "customer": { "tier": "pro" } }, "current": { "action": "approve" }, "source": "jev", "confident": true } ] } ``` - `state` (string | object, required): The real state, as your program sends it. - `current` (object, required): The decision you made, per question. - `source` (string): Who made it: `jev` (the default), `process` (your own code or process) or `person`. - `confident` (boolean): Optional: the current decision was a confident one. Your model decides each one with its serving version and your hard rules. Nothing is acted on, nothing is billed as a decision, and nothing goes to your decision log. ## Review the disagreements **`GET /v1/domains/{domain}/shadow/review`** **`POST /v1/domains/{domain}/shadow/review`** The review list is short (at most 20 at a time), your model's most confident disagreements first. For each, say which answer was right (`"pick": "model"` or `"current"`), give the right one (`"answers"`), or `"skip"`. The same list is in the console, under the model's review queue. `GET /v1/domains/{domain}/shadow` shows how often they agree. ## What your model learns from it Not every answer teaches the same: 1. **Real outcomes** (what really happened) teach most. 2. **A person's review** comes next: your shadow reviews, the review queue, spot checks. 3. **A confident Jev answer** teaches least, and only when Jev was confident. Your next retrain uses them all, weighted this way. Everything else in shadow mode only measures. ## Spot checks Once your model acts, a few of its confident decisions are set aside at random for a person to check: `light` (about 1 in 100, the default), `thorough` (about 2 in 100) or `off`. Nothing is held up and nothing pops up: they wait on one optional weekly card in the model's review queue ("Check these 10 decisions, about 2 minutes"). **`GET /v1/domains/{domain}/spot-checks`** Answer each like the review queue (`POST /v1/domains/{domain}/answers`). Progress shows one line: "spot-checked accuracy: 94%". Change the setting in the console's settings, with `PATCH /v1/settings {"spot_checks": "thorough"}`, or `canonopy spot-checks --set thorough`. --- Source: https://canonopylabs.com/docs/training-and-reports (Markdown: https://canonopylabs.com/docs/training-and-reports.md) # Training and reports Everything that makes a model runs on our servers. You send cases and get back a version with a report. A new version usually takes about a minute. The quickest start needs no past cases: build a model from the JSON you send Jev, see [Start with no data](https://canonopylabs.com/docs/start-with-no-data.md). This page covers the other way in, describing the decision to the set-up agent, and what every version's report tells you. A model made that way uses your past cases as described in [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md), and its retrains are [targeted retrains](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain). ## Set up the domain **`POST /v1/domains`** Most people start in the console: [New decision model](https://canonopylabs.com/console/domains/new) opens a conversation with the set-up agent. The same thing over the API: **`POST /v1/domains/{domain}/agent`** ```json { "message": "We handle refund requests. Actions: approve, escalate, decline. Fields: message (text), amount (number), tier (category: free, pro, enterprise), fraud_flag (yes/no). Never approve when amount is over 500. Always escalate when fraud_flag is true." } ``` The reply tells you what it set up, in plain words, and what to do next (`next_step`: `describe` while it still has questions for you, then `train`). For example: *"Set up 3 actions, 4 state fields and 2 rules that can't be broken. Next, train it (looking at the example cases is optional, any time)."* You can [train](#train) right away. Uploading past cases first is optional, and so is reviewing the examples. ## Upload history **`POST /v1/domains/{domain}/data`** Each past case and how it was resolved: the outcome you already record is enough. Send a `multipart/form-data` file (`.csv` or `.jsonl`), a `text/csv` or `application/x-ndjson` body, or JSON `{"rows": [...]}`. A row is either `{"state": {...}, "answers": {"question": value}}` or a flat row whose columns are state fields and question ids. ```bash curl https://api.canonopylabs.com/v1/domains/refunds/data \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -F file=@history.csv ``` ```json { "accepted": 9800, "rejected": 12, "total_cases": 9800, "per_question": { "action": 9800 }, "problems": [{ "row": 17, "problem": "action: 'refnd' is not a valid answer (options: approve, escalate, decline)" }] } ``` Every row is checked. A row with a number that isn't a number, an answer that isn't one of the options, or a question the model doesn't have is skipped, counted in `rejected` and listed with its row number (the first 20). Empty values, `NaN` and infinity count as missing. Uploads can be up to 50 MB; split bigger files. **Your answers are the truth.** Where your history says how a case was resolved, that is what your model learns and is measured against. Your rules still can't be broken: a past answer that one of them doesn't allow is counted in the report (`fix_history`). On a model made with `POST /v1/build`, uploaded cases are used as in [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md): rules found in them are proposed for you to confirm, and cases that contradict your rules wait for your review. Each upload is kept as one unit with an id (`upload_id` in the answer), so you can replace or remove it later. Add `?file_name=history.csv` to name a raw CSV or JSONL body in the upload list; a multipart upload uses the file's own name. ## Replace or remove cases Improve the model you already have instead of starting a new one. Your code keeps calling `refunds@latest`; the next training run makes `refunds@2`, `refunds@3`, and so on, from the cases the model has at that moment. **`GET /v1/domains/{domain}/uploads`** Every upload, newest first: its id, date, file name, how many cases it holds, and how many rows were set aside. Examples you corrected are listed too (`source: "signoff"`). Cases uploaded before uploads had ids appear as one group with the id `earlier`. ```json { "domain": "refunds", "total_cases": 6361, "from_uploads": 6349, "from_signoff": 2, "from_outcomes": 12, "uploads": [ { "id": "upl_4f1c2a9b0d3e5f67", "source": "upload", "mode": "replace", "file_name": "cases-sep.csv", "created_at": "2026-09-27T09:12:00Z", "cases": 6347, "rejected": 3 }, { "id": "upl_0a1b2c3d4e5f6a7b", "source": "signoff", "mode": "add", "file_name": null, "created_at": "2026-09-20T15:02:00Z", "cases": 2, "rejected": 0 } ] } ``` **Replace every uploaded case** with a new file, in one step: ```bash curl "https://api.canonopylabs.com/v1/domains/refunds/data?mode=replace" \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -F file=@better-history.csv ``` The earlier uploaded cases are deleted once the new file's rows are in, and `replaced` says how many went. If no row of the new file is accepted, nothing is stored and nothing is removed. **Delete one upload:** **`DELETE /v1/domains/{domain}/uploads/{upload_id}`** **Delete every uploaded case.** So it can't happen by accident, confirm with the model's name: **`DELETE /v1/domains/{domain}/data?confirm={domain}`** ```json { "domain": "refunds", "removed": { "uploaded_cases": 6278, "signoff_cases": 0, "decisions": 0 }, "uploads_removed": ["upl_9e8d7c6b5a4f3e2d"], "remaining": { "total_cases": 14, "from_uploads": 0, "from_signoff": 2, "from_outcomes": 12 }, "message": "Deleted permanently: 6,278 uploaded cases. Train again to make the next version from the cases left. Versions already trained are kept: they're yours, and still download and serve." } ``` What each call removes: | call | removes | keeps | |---|---|---| | `POST …/data?mode=replace` | every uploaded case, once the new file's rows are in | examples you corrected, outcomes | | `DELETE …/uploads/{upload_id}` | that upload's cases (or a review's corrected examples, when you name them) | everything else | | `DELETE …/data?confirm=…` | every uploaded case | examples you corrected, outcomes | | `…&signoff=true` | also the examples you corrected | the review record: who reviewed, when, and which examples | | `…&outcomes=true` | also the hosted decisions the next version would learn from: those with a reported outcome or an unsure-queue answer, and those our hosted backup answered | other hosted decisions | | `…&starter=true`, or `DELETE …/uploads/starter` | also the [starter cases](https://canonopylabs.com/docs/day-one.md#a-starter-model-for-decisions-on-text) of a question that reads text (they then stay off for this model) | everything else | - **Deleting is permanent.** The cases are removed from our database, not hidden. Backups roll off on their usual schedule. - **Versions already trained stay yours.** They keep serving and downloading; roll back to one any time. - **Examples you corrected stay unless you remove them.** They are answers you gave yourself, so replacing or clearing uploads leaves them in place. - **While a new version is being made**, replacing and deleting answer `409`: wait for the job (`GET /v1/jobs/{id}`), then try again. Adding cases is always fine. Then train again. The report's `cases` says what changed ("Trained on 6,349 cases; 6,278 earlier cases were removed since version 3"), and the comparison with the serving version uses the current held-back cases with their current answers. So new cases that fix a bad habit can pass it, and a version that is really worse on your current cases is still kept out. The console does the same on the model's **Data** tab: the upload list with a delete button per upload, **Replace all cases** when you upload a file, and **Clear all** after you type the model's name. ## Train **`POST /v1/domains/{domain}/train`** ```json { "promote": "auto", "wait": true } ``` - `promote: "auto"` promotes the new version when its report recommends it; `always` and `never` override that. - With `wait: false` you get a job to poll at `GET /v1/jobs/{id}`. - It learns from the cases the model has now. After a replace or a delete, that means the current cases only. The answer is a job with the report attached: ```json { "id": "job_0a1b2c3d4e5f6a7b", "domain": "bank-support", "kind": "train", "status": "done", "version": 3, "message": "Version 3 is serving.", "report": { "…": "…" } } ``` ## Review the examples (optional) **`GET /v1/domains/{domain}/examples`** **`POST /v1/domains/{domain}/signoff`** About 20 example cases, each with the answer your setup gives and why. Look at them any time, before or after training: nothing waits on them. If you have history, [upload it](#upload-history) first: the examples are then drawn from your own cases. To correct one, send only the ones you correct. `"retrain": true` starts the retrain in the same call, and the answer has its `job_id` and a `message`: ```json { "examples": [ { "id": "ex_02", "ok": false, "correct": { "action": "escalate" }, "note": "pro accounts under 30 days go to a person" } ], "retrain": true } ``` Without `retrain`, your corrections are learned at the next retrain. The review is kept as a record. In the console, it's **Review examples** on the model's page; in the CLI, `canonopy examples refunds` and `canonopy signoff refunds --correct ex_02:action=escalate --retrain`. ## The report **`GET /v1/domains/{domain}/versions/{version}/report`** Every number comes from **your own held-back cases**, which are never used for learning. The one exception is a [starter model](https://canonopylabs.com/docs/day-one.md#a-starter-model-for-decisions-on-text) for a question that reads text, before you have enough cases of your own: its accuracy is measured on starter example cases, marked `measured_on: "starter_cases"` and explained in the report's `starter` section. A model you [built with no data](https://canonopylabs.com/docs/start-with-no-data.md) is the same: until you have enough answered cases of your own (about 200), it is measured on held-back examples of situations, marked `measured_on: "situations"`. From then on those examples only teach (at full weight, they never fade), and the accuracy, the calibration and the never-worse check use your own held-back cases only; the gate's `measured_on` says which. The console shows it as cards; the JSON has four parts. ### 1. Summary - accuracy per question, against the previous version and your day-one backend (`previous`, `day_one`, `change_points`); - how it did on rare situations it was tested on; - calibration per question, with a plain sentence ("When it says 90% sure, it's right 89% of the time"); - how many cases each rule matched and how many decisions it changed. Violations are always 0; - a `promote` or `keep` recommendation, with the reason. `cases` says how many of your cases the version learned from and, when cases were deleted since the last version, how many. `starter` (when a question reads text and has few of your own cases) says whether a starter model is answering, still helping some options, or has stepped aside. ### 2. Weak spots - the weakest options, with their accuracy and case counts; - the most-confused pairs, with real example cases; - for day-one questions: how many more outcomes until they graduate. ### 3. What would make it better A ranked list, each item with an estimated gain and a `fix`: the exact API call that acts on it, so the console can offer it as one click. | `kind` | What it suggests | |---|---| | `more_cases` | more cases where accuracy is still climbing, estimated from your own learning curve | | `clarify_rule` | a boundary your rules don't separate: say once which way the example cases go | | `add_option` | an option for cases that fit none | | `more_language_cases` | a language below the bar: "upload ~300 cases in Swahili to handle more automatically" | | `fix_history` | past cases whose answer one of your rules doesn't allow: your rule was followed for them; change the rule if it's too strict | ### 4. The safety bar The bar this version set for itself, in plain words: "automatic answers 97%+ accurate; 78% handled automatically; the rest double-checked". For each question: the confidence an answer needs, the share of your held-back cases handled automatically, and how often those were right. Then per language: the strong languages (answered automatically) and the ones still double-checked. There's no threshold to set; each retrain measures it again. See [Languages](https://canonopylabs.com/docs/languages.md). ## Versions **`GET /v1/domains/{domain}/versions`** **`POST /v1/domains/{domain}/promote`** `@latest` serves whichever version you promoted. A roll-back is a promote of an older version: `{"version": 2}`. To see how accuracy moved from version to version, and what the next retrain would learn from, open the model's **Progress** tab or call `GET /v1/domains/{domain}/progress`. See [See what's waiting, then retrain](https://canonopylabs.com/docs/improving.md#see-whats-waiting-then-retrain). --- Source: https://canonopylabs.com/docs/improving (Markdown: https://canonopylabs.com/docs/improving.md) # Improving Each of these is something you do once; the next version takes about a minute. Every training report ranks them for you. For a model made from a Jev request ([`POST /v1/build`](https://canonopylabs.com/docs/start-with-no-data.md)), also see [Improve your model](https://canonopylabs.com/docs/improve-your-model.md): advice on what would help most, targeted retrains with a before and after, rule changes that retrain straight away and say what they change, and marking a decision wrong (which works for every model). ## See what's waiting, then retrain **`GET /v1/domains/{domain}/progress`** Your model learns from new cases only when you retrain it: nothing retrains on its own, and nothing is sent to you. The console's **Progress** tab (and `canonopy progress refunds`) shows what the next retrain would learn from, since the latest version: - uploaded past cases, reported outcomes, unsure cases you (or our backup) answered, cases synced from offline devices, examples you corrected, and [decisions you marked wrong](https://canonopylabs.com/docs/improve-your-model.md#mark-a-decision-wrong); - how many of the new outcomes differ from the answer it gave: those teach it most. It says it in one line: "48 new cases since version 4, 12 where the real outcome differed from its answer. Retrain to learn from them." Then retrain with the button, `canonopy train refunds` or `POST /v1/domains/{domain}/train`. The counts start again from the new version. It also shows [advice](https://canonopylabs.com/docs/improve-your-model.md#advice): what would improve your model most right now. The same view lists every version: when it was made, how many of your cases it learned from, its accuracy per question and the change against the version before. When the report compared both versions on the same held-back cases, that is the number shown; otherwise each version was measured on its own held-back cases, and it says so. A starter model is measured on starter example cases, not your own, and is marked as such. Once a version has 20 reported outcomes, you also see how often its answers matched them. ## Report outcomes **`POST /v1/outcomes`** Tell us what really happened on a decision. Outcomes feed the next version and day-one graduation. ```json { "decision_id": "dec_4f1c2a9b0e7d6a51", "action": "block-card", "answers": { "urgency": 2, "needs_human": false } } ``` ```json { "recorded": true, "decision_id": "dec_4f1c2a9b0e7d6a51", "graduation": { "urgency": { "status": "fallback", "version": null, "outcomes": 213, "needed": 787 } } } ``` `action` is shorthand for the decision question's answer. ## The unsure queue **`GET /v1/domains/{domain}/unsure?limit=50`** **`POST /v1/domains/{domain}/answers`** Decisions held for a person wait here until someone answers them, in the console or over the API: `review` (still below the model's safety bar after the message was double-checked) and `ask` (your rules allow no action). ```json { "decision_id": "dec_7a…", "answers": { "action": "route-payments" } } ``` Every answer is captured, and the next version learns from it, so it keeps getting better on your own traffic. Your own LLM or Jev can answer the queue too: see [Using with LLMs](https://canonopylabs.com/docs/using-with-llms.md). ## Add an option or an action **`POST /v1/domains/{domain}/options`** A new Choice option or action in about 30 seconds, without breaking the others: ```json { "question": "action", "option": "freeze-account", "description": "freeze the whole account, including the app and all cards", "when": "topic is lost_stolen or unrecognised and new_device_login_24h is true", "never_when": [{ "field": "new_device_login_24h", "op": "is_false" }] } ``` The job's report has a before/after: ```json { "question": "action", "option": "freeze-account", "took": "under a minute", "before": { "accuracy": 0.958, "cases": 988 }, "after": { "accuracy": 0.955, "cases": 988, "option_recall": 0.93, "option_cases": 58, "unchanged_cases_accuracy": 0.957 } } ``` In the bank test, adding `freeze-account` took 29 seconds and left the other nine actions at 96.0%. ## Clarify a rule where it hesitates Most mistakes sit on blurry boundaries, like "is this a card question or a payments question?". The report shows those cases; say once, in plain words, which way each goes (`POST /v1/domains/{domain}/agent`), and train again. On a model made with `POST /v1/build`, a rule change shows what it changes first: see [Change a rule with a preview](https://canonopylabs.com/docs/improve-your-model.md#change-a-rule-with-a-preview). ## Upload more of what's weak If lost or stolen cards are 89% right and climbing with data, the report estimates what 500 more such cases would add. Upload them and train. ## Replace cases that taught a bad habit If some of your past cases were resolved badly (old recordings, a policy you've since changed), don't start a new model: replace them and train the same one. Your code keeps calling `refunds@latest`, and the next version is `refunds@2`. ```bash curl "https://api.canonopylabs.com/v1/domains/refunds/data?mode=replace" \ -H "Authorization: Bearer $CANONOPY_API_KEY" -F file=@better-history.csv curl -X POST https://api.canonopylabs.com/v1/domains/refunds/train \ -H "Authorization: Bearer $CANONOPY_API_KEY" -H "Content-Type: application/json" -d '{"wait": true}' ``` The new version is compared with the serving one on your current held-back cases, with their current answers, so better cases can win. To remove just one upload instead, list them (`GET /v1/domains/{domain}/uploads`) and delete it. Deleting is permanent; versions already trained are kept. See [Replace or remove cases](https://canonopylabs.com/docs/training-and-reports.md#replace-or-remove-cases). --- Source: https://canonopylabs.com/docs/running-it-yourself (Markdown: https://canonopylabs.com/docs/running-it-yourself.md) # Running it yourself Every version you train is yours to keep: download it any time, during the trial, while you subscribe, or after. Run it on your own servers, a laptop, or offline, with the same request and answer format as the hosted endpoint. **`GET /v1/models/{domain}@{version}/download`** ```bash curl -L -o bank-support@3.zip \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ https://api.canonopylabs.com/v1/models/bank-support@3/download ``` `@latest` works too. The zip holds three files, plus a short `README.txt` on how to run it: | File | What's in it | |---|---| | `model.json` | runtime settings: questions, options, state fields, the text reader it uses, and its safety bar | | `heads.onnx` | the trained model, in the open ONNX format (the runtime runs it with ONNX Runtime) | | `rules.json` | your rules, as a table of allowed and blocked actions | ## The runtime The runtime is the one package you install to run a model yourself; using the hosted API needs nothing installed. It's served from this site: ```bash # models that read text (about 100 languages) pip install "canonopy-runtime[text] @ https://canonopylabs.com/dl/canonopy_runtime-0.6.3-py3-none-any.whl" # numbers and categories only pip install https://canonopylabs.com/dl/canonopy_runtime-0.6.3-py3-none-any.whl ``` ```python from canonopy_runtime import load model = load("bank-support@3.zip") res = model.decide({"state": {"message": "I lost my card", "amount": 40, "fraud_flag": False}}) print(res["decision"]) # {'question': 'action', 'action': 'block-card', 'confidence': 0.98, 'blocked_by_rules': [], ...} ``` From the command line: ```bash python -m canonopy_runtime bank-support@3.zip request.json ``` Your rules are enforced on every decision, exactly as on the hosted endpoint. Models downloaded before September 26 (with `weights.npz` instead of `heads.onnx`) keep working with the same runtime. ## Unsure messages, offline The runtime runs the same two checks as the hosted endpoint: a message in a language the model isn't strong in, or below its safety bar, is double-checked. With `CANONOPY_API_KEY` set, it hands that message to us (translated to English and decided again, then your day-one backend). Offline, or when we can't be reached, it goes to **your own fallback** if you gave one; otherwise it is held for review (`"route": "review"`). It never guesses, unless you turn on best guess (below). `load(path, service=None)` never calls out. See [Languages](https://canonopylabs.com/docs/languages.md). ### Best guess, for games For games, turn on best guess: the model always acts, and tells you how sure it was. ```python model = load("snake@1.zip", best_guess=True) # or model.decide(request, best_guess=True) for one call res = model.decide({"state": state}) res["decision"] # {'action': 'right', 'confidence': 0.62, ..., 'route': 'act', 'best_guess': True} ``` In a game, "unsure" usually means two moves are about as good, not a real doubt. With best guess on, an answer below the safety bar acts on the model's most likely answer instead of being routed, and is marked `"best_guess": true` (on the answer and on the decision), with its confidence and probabilities as usual. A confident answer has no `best_guess` field, so you can always tell the two apart. - **Your rules still apply first**: a blocked option is never the guess. When your rules allow no option at all, the decision still says `"action": null`, `"route": "ask"`. - Yes/no and score questions work the same way: the most likely answer, marked. - A best guess doesn't call us, doesn't call your fallback, and isn't saved for sync. - A message in a language the model isn't strong in is still routed as above: best guess covers "unsure", never "can't read this language". - **Off by default**: business decisions keep routing unsure cases to a person. `model.decide(request, best_guess=False)` turns it off for one call. ### Your own fallback Pass a function, and the runtime calls it for the questions it would have handed to us, whenever we can't be reached (or no key is set). When we can be reached, we answer first and your function isn't called. ```python from canonopy_runtime import load def my_fallback(state, questions): """state: the case. questions: {question id: its definition}, only the ones that need a second look. Return {question id: answer}, or None to hold them for review.""" if "action" in questions and state.get("amount", 0) > 500: return {"action": "escalate"} # your own rule of thumb return None model = load("refunds@3.zip", offline_fallback=my_fallback) res = model.decide({"state": {"message": "Mi pedido llegó roto", "amount": 900}}) res["answers"]["action"] # {'type': 'choice', 'choice': 'escalate', 'confidence': 1.0, ..., 'source': 'local_fallback', 'route': 'act'} ``` - **An answer** is the usual shape (`{"type": "choice", "choice": ..., "confidence": ..., "probabilities": {...}}`, a Score's `probabilities` or `score`, a Noul's `noul`), or just the value: an option label, a level number, or `true`/`false` (confidence 1). - It comes back marked `"source": "local_fallback"`. **Your rules still apply after it**: an action your rules block is never returned. - **None, an error, or an answer that isn't one of the options** means `"route": "review"`, as without a fallback. - One call can use its own: `model.decide(request, fallback=other_function)`. A small local model works the same way: ```python def small_model(state, questions): if "topic" not in questions: return None label, p = my_classifier.predict(state["message"]) # any model you run yourself return {"topic": {"type": "choice", "choice": label, "confidence": p}} if p >= 0.9 else None ``` ## Save and sync later Every case that needed a second look while we couldn't be reached is **saved on your machine**, whether your fallback answered it or not, so a person can answer it later and your next version learns from it. The response then carries an `id`. - **Where:** `offline-cases.jsonl` (one JSON object per line: the case, the questions, the answers and their `source`, the time, the model version and the `id`), in `CANONOPY_OFFLINE_DIR`, else your user data folder: `~/Library/Application Support/canonopy/offline` on macOS, `%LOCALAPPDATA%\canonopy\offline` on Windows, `~/.local/share/canonopy/offline` on Linux. It's your file: read it, copy it or delete it any time. - **Bounded:** the newest 10,000 cases or 50 MB, whichever comes first; past that, the oldest go. - **Off:** `load(path, queue=None)` (or `CANONOPY_OFFLINE_QUEUE=off`) saves nothing; `queue="some/folder"` picks the folder. **Sync** when you're back online: ```python res = model.sync() # {'sent': 12, 'accepted': 12, 'already_there': 0, 'rejected': [], 'remaining': 0, 'decision_ids': {...}, ...} ``` ```bash canonopy-runtime sync # or: canonopy sync (reads CANONOPY_API_KEY) canonopy-runtime queue # where the file is, and how many cases wait ``` It also happens by itself, in the background, after the next call that reaches us (`load(..., auto_sync=False)` turns that off). - Cases are sent in batches to `POST /v1/domains/{domain}/offline-cases` and **leave the file only once they are stored**. Sending one twice is harmless. - They join your model as normal cases: the ones still held wait in the **Unsure** view, marked *from an offline device*. Our hosted backup may answer some of them first; the rest wait for you. Every answer is learned by the next version. - **Report outcomes** on them with the same `id`, once synced: `POST /v1/outcomes {"decision_id": "loc_…", "action": "refund"}`. - Syncing needs an active subscription or trial, like hosted decisions. After the trial it's refused (`402`) and **your cases stay in the file** until you sync again. ## Size and speed A decision model is well under 1 MB; models that read text also use a shared text reader (`reader-en-1`, about 34 MB, or `reader-multi-1`, about 113 MB; fetched once and shared by every model; set `CANONOPY_READER_PATH` to a folder holding it to run fully offline). Their licences and attribution are in the `THIRD_PARTY_NOTICES` file that comes with the runtime. In the bank benchmark, on a laptop CPU, a decision took about 0.5 ms once the text was read and about 39 ms including reading the text. To free the space, `canonopy-runtime reader remove` (runtime 0.6.2 and later) deletes every downloaded reader; a model that reads text downloads its reader again the next time it needs it. `canonopy-runtime reader list` shows what's there. ## Godot Godot 4 games use the Canonopy addon instead: no Python, no ONNX Runtime, the same decisions, from GDScript. See [Use your model in Godot](https://canonopylabs.com/docs/godot.md). ## Phones and other runtimes (coming soon) The runtime is Python and Godot today. Other runtimes follow. --- Source: https://canonopylabs.com/docs/coding-agents (Markdown: https://canonopylabs.com/docs/coding-agents.md) # Use with coding agents Coding agents like Claude Code, Cursor and Codex can use Canonopy Decisions directly: build your own model from the JSON your code sends Jev, switch the code over once it's ready, show you how it decides if you ask, and improve it later, without you writing the calls. Connect them to the hosted MCP connector, which needs nothing installed, add the [skill](#skill) so they know when and how to use it, or point them at these docs. ## The hosted MCP connector The connector runs on our servers at `https://api.canonopylabs.com/mcp` (MCP Streamable HTTP). There's nothing to install: your agent needs an API key from the [console](https://canonopylabs.com/console/keys), sent in an `Authorization` header like any API call. ### Claude Code ```bash claude mcp add --transport http canonopy https://api.canonopylabs.com/mcp --header "Authorization: Bearer cnp_…" ``` Add `--scope user` to have it in every project. Check it with `claude mcp list`, or `/mcp` inside a session. ### Cursor `.cursor/mcp.json` in your project, or `~/.cursor/mcp.json` for every project: ```json { "mcpServers": { "canonopy": { "url": "https://api.canonopylabs.com/mcp", "headers": { "Authorization": "Bearer cnp_…" } } } } ``` ### Codex `~/.codex/config.toml`, with your key in the `CANONOPY_API_KEY` environment variable: ```toml [mcp_servers.canonopy] url = "https://api.canonopylabs.com/mcp" bearer_token_env_var = "CANONOPY_API_KEY" ``` ### Any other MCP client Add a Streamable HTTP server with the URL `https://api.canonopylabs.com/mcp` and the header `Authorization: Bearer cnp_…`. It's JSON-RPC 2.0 over `POST`, answered with JSON; there's no session to keep and no stream to open (`GET` answers `405`). To see it work with `curl`: ```bash curl https://api.canonopylabs.com/mcp \ -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -H "MCP-Protocol-Version: 2025-06-18" \ -d '{"jsonrpc": "2.0", "id": 1, "method": "tools/list"}' ``` A missing or refused key gets `401` with a JSON error. Protocol versions `2025-06-18`, `2025-03-26` and `2024-11-05` are supported. ### Every tool call is an API call Each tool call is a normal call to the API with your key: the same fair-use limit (50 requests per second per key), the same errors, the same plan. Nothing is billed per tool call. ### What the hosted connector doesn't do It runs on our servers, so it never reads or writes files, and never reads environment variables: - **`upload_history`** and **`replace_history`** take past cases inline: `rows` (a list of cases), or a file's contents as `text` (CSV with a header row, JSONL, or a JSON list), up to about 2 MB per call. Your agent reads the file and sends its text; split bigger files, or use `canonopy upload`. - **`build_model`** takes past cases inline as `history` rows; reading them from a file on disk (`history_file`) is the local server's. - **`download_model`** returns a link that works for 10 minutes without a key, and the `curl` line to save it. - **`set_day_one_fallback`** takes no keys: add your Jev or OpenAI-compatible key once in the console (open the decision model, then **Day one**). `local` and `none` need no key. For those on your own machine, use the local MCP server. ## The local MCP server `canonopy mcp` runs the same tools on your machine, over stdio, from the `canonopy-cli` package. Use it when you want the agent to upload files straight from disk, save downloaded models to disk for [offline use](https://canonopylabs.com/docs/running-it-yourself.md), or read a day-one backend's key from an environment variable. It needs Python 3.9 or later and nothing else. ```bash pip install https://canonopylabs.com/dl/canonopy_cli-0.6.2-py3-none-any.whl ``` The key goes in the server's environment, never in the chat. ### Claude Code ```bash claude mcp add canonopy --env CANONOPY_API_KEY=cnp_… -- canonopy mcp ``` ### Cursor ```json { "mcpServers": { "canonopy": { "command": "canonopy", "args": ["mcp"], "env": { "CANONOPY_API_KEY": "cnp_…" } } } } ``` ### Codex ```toml [mcp_servers.canonopy] command = "canonopy" args = ["mcp"] env = { CANONOPY_API_KEY = "cnp_…" } ``` ### Any other MCP client Run `canonopy mcp` as a stdio server (JSON-RPC 2.0, one message per line). If `canonopy` isn't on the client's `PATH`, give the full path (`which canonopy`) or run `python -m canonopy_cli mcp`. - `CANONOPY_API_KEY` (environment variable, required): Your API key. Sent only to the API; no tool ever prints or returns it. - `CANONOPY_BASE_URL` (environment variable): The API. Default `https://api.canonopylabs.com`. - `CANONOPY_DOCS_URL` (environment variable): The docs the `docs` tool reads. Default `https://canonopylabs.com`. ## Tools The hosted connector and the local server have the same 47 tools, in the order you'd usually use them. A decision model is called a *domain* in the API; the tools say decision model. **From the JSON you send Jev** ([Start with no data](https://canonopylabs.com/docs/start-with-no-data.md)): - `build_model`: your own model from a Jev request: `state` (one real example), `questions` (as you send them to Jev), and optionally `rules` in plain words, `name`, `description` and `history` rows (your past cases). The local server also takes `history_file`, a `.csv`, `.jsonl` or `.json` file on disk. **JSON first:** without `state` and `questions` it builds nothing and says how to get them (ask you to paste the request you send Jev, or draft it from your code and confirm it with you). It asks [a few questions about the JSON](https://canonopylabs.com/docs/answer-a-few-questions.md) first by default (`interview`, `answer_timeout_seconds`); it returns when the model is ready, or sooner if those questions wait for your answers. See [Build with your coding agent](https://canonopylabs.com/docs/build-with-your-coding-agent.md). - `draft_request_help`: the JSON's shape, an example and a checklist for getting it. - `get_build_questions`, `answer_build_questions`: the build's questions about your JSON (`build_id` or `model`), and your answers (`answers`: each `id` with the keys its `answer_with` lists, or `skip`; `done`). - `get_build`: where a build is (`build_id`, or `model` for its latest build): the phases, what it read, the quality check, the examples once it's ready, rules found and flagged cases. With `wait: true` it returns once the build is ready (or needs attention, failed, or waits for answers); if the answer has `still_working`, call it again. Nothing is needed from you meanwhile. - `confirm_found_rules`: confirm or reject the [rules found in your history](https://canonopylabs.com/docs/rules-found.md) (`build_id` or `model`; `rules`: each `id`, `decision` and `hard`). - `review_flagged_history`: past cases that contradict your rules (`build_id` or `model`). With no decisions it lists the pending ones; `rows` (each `id` and `decision`: `follow_rule`, `keep` or `drop`) or `all` decide them. - `get_advice`: what would [improve the model](https://canonopylabs.com/docs/improve-your-model.md) most, in plain words (`model`). - `change_rules`: new rules in plain words (`model`, `rules`, `mode`: `add` or `replace`). It retrains straight away, and the build says what changed. - `mark_decision_wrong`: the right answer for a decision (`decision_id`, `answers` or `action`, `note`); the next retrain learns it. - `shadow_mode`, `send_shadow_cases`, `review_shadow_disagreements`: [shadow mode](https://canonopylabs.com/docs/shadow-mode.md): your model beside your current decisions on real requests, without acting, and the short review list. - `report_play_results`: a [playtest](https://canonopylabs.com/docs/playtest-your-model.md)'s results (`canonopy playtest` sends them for you). - `spot_checks`: this week's optional spot-check card (`model`), or the setting (`set`: `off`, `light`, `thorough`). On a built model, `train` starts a targeted retrain and returns its build id. `get_examples` shows how it decides on about 20 situations, and `correct_examples` retrains with your corrections: both optional. **Every model:** - `list_decision_models`, `get_decision_model`: what exists, and where each stands (status, actions, rules, fields, serving version). - `create_decision_model`: start one, with a name and optionally a plain-words description. - `describe_decision_model`: the set-up conversation (describe the decision, answer open questions, refine it). - `update_decision_model`: exact changes to questions, rules or fields; archive or restore. - `get_examples`, `correct_examples`: optional, any time: about 20 example cases showing how it decides, and your corrections (`corrections`, and `retrain`, true by default, to retrain with them now). The old name `sign_off_examples` still works. - `upload_history`, `history_summary`: past cases and how each was resolved: inline rows or a file's text (hosted), or a local `.csv`, `.jsonl` or `.json` file (local server). - `list_uploads`, `delete_upload`, `replace_history`, `clear_history`: see each upload, delete one, swap every uploaded case for new ones, or delete them all, then `train` for the next version of the same model. These delete permanently, so each needs `confirm` set to the decision model's name. - `train`, `get_job`: a new version in about a minute, with its report. - `get_report`, `list_versions`, `promote_version`: read reports, see versions, promote or roll back. - `get_progress`: what's waiting for the next retrain, every version's accuracy, and advice. - `decide`: a decision in Jev's format, with the rules applied and `act` / `review` / `ask`. - `report_outcome`: what really happened, for the next version. - `list_unsure`, `answer_unsure`: the unsure queue. - `list_decisions`, `get_decision`, `export_decisions`: the [decision log](https://canonopylabs.com/docs/decisions-and-rules.md#the-decision-log): past decisions with filters and search, one in full, or a CSV or JSONL file (saved to disk by the local server; a short-lived link from the hosted connector). - `add_option`: a new option or action, with a before/after report. - `get_day_one_fallback`, `set_day_one_fallback`: the backend for questions not learned yet. - `usage`: plan, decisions this month and the monthly price. - `download_model`: a finished version, to [run yourself](https://canonopylabs.com/docs/running-it-yourself.md): a short-lived link (hosted) or a zip saved to disk (local server). - `docs`: these docs (the index, one page, a search, or one API endpoint explained), and the [skill](#skill) with `page=skill`. Errors come back as readable tool errors with the API's message, like `409 conflict: A domain with that name already exists`, so the agent can fix the call. ## Skill The Canonopy Decisions skill is a packaged instruction file your agent loads when a task calls for it: when a decision model fits and when it doesn't, every step of the workflow with its tool and API call, how to read `source`, `route` and `decision`, switching from Jev, pricing, and what to do about each error. It's plain Markdown: `SKILL.md`, plus `reference/` and `examples/` files the agent reads only when it needs them. ### Claude Code With the [CLI](https://canonopylabs.com/docs/cli-reference.md) (`canonopy-cli` 0.3.1 and later): ```bash canonopy skill install # for you, in every project: ~/.claude/skills/canonopy-decisions/ canonopy skill install --project # this project only: .claude/skills/canonopy-decisions/ ``` It prints where it went, and never replaces an installed copy unless you add `--force`. Without the CLI, unzip the skill into your skills folder: ```bash curl -fsSL -o canonopy-decisions-skill.zip https://canonopylabs.com/dl/canonopy-decisions-skill.zip unzip canonopy-decisions-skill.zip -d ~/.claude/skills/ ``` Claude Code picks it up in new sessions, and uses it whenever you ask about repeated decisions, switching from Jev, or Canonopy itself. ### Other agents - **The files:** [canonopylabs.com/skill/SKILL.md](https://canonopylabs.com/skill/SKILL.md), with its supporting files next to it (for example `/skill/reference/troubleshooting.md`), or the whole `canonopy-decisions/` folder as a [zip](https://canonopylabs.com/dl/canonopy-decisions-skill.zip). Put the folder where your agent reads skills (`canonopy skill install --dir PATH` copies it there), or tell it to read `SKILL.md` first. - **Through MCP:** the hosted connector and the local server both offer it as the resource `canonopy://skill/SKILL.md` (its supporting files are resources too), and their instructions point agents to it. The `docs` tool returns it with `page=skill`, for clients that don't read resources. ## What to ask - "Build a Canonopy model from the Jev request in `src/refunds.py`, with our refund rules, and switch the code over once it's ready." - "Here are our past refunds in `data/refunds.csv`: rebuild `refunds` with them and walk me through the rules it found." - "What does the advice for `refunds` say? Retrain it if it's a weak spot, and show me the before and after." - "Set up a decision model for refund requests: approve, escalate or decline. Never approve over $500; fraud-flagged accounts always go to a person." - "Show me how `refunds` decides on the example cases, correct the ones I say are wrong, and retrain." - "Upload `data/past_refunds.csv` to refunds and train it. What does the report say would help most?" - "Decide this case with `refunds@latest` and explain the answer." - "Go through the unsure queue for refunds with me." - "How do I switch our Jev code over?" ## Keys stay out of the chat - Your API key lives in the connector's configuration (the `Authorization` header) or the local server's environment. It's sent only to the API, and no tool returns it. - Your day-one backend's key (a Jev or OpenAI-compatible key) goes in the console. With the local server, it can instead come from an environment variable you name: add it to the server, for example `--env TYPESAFE_API_KEY=…`, and ask the agent to "use TYPESAFE_API_KEY". - The tools that change what production answers (`correct_examples`, `promote_version`, `answer_unsure`) tell the agent to check with you first. So does the [skill](#skill) for those that change what your model learns: `confirm_found_rules`, `review_flagged_history`, `change_rules` and `mark_decision_wrong`. The tools that delete cases (`delete_upload`, `replace_history`, `clear_history`) say they delete permanently and won't run until the agent passes the decision model's name as `confirm`. ## Point your agent at the docs - **[llms.txt](https://canonopylabs.com/llms.txt)**: an index of every docs page as Markdown, with a line about each. - **[llms-full.txt](https://canonopylabs.com/llms-full.txt)**: every docs page in one file. - **[The skill](https://canonopylabs.com/skill/SKILL.md)**: when and how to use Canonopy Decisions, step by step (see [Skill](#skill)). - **Markdown pages:** add `.md` to any docs URL, like `https://canonopylabs.com/docs/quick-start.md`. The introduction is `/docs.md`. - **[The API description](https://api.canonopylabs.com/openapi.json)**: every endpoint and field, with examples and a "start here" workflow. Browse it at [api.canonopylabs.com/docs](https://api.canonopylabs.com/docs). A prompt that works: *"Read https://canonopylabs.com/llms.txt, then call our `refunds` decision model before a refund is issued, and hand `ask` decisions to a person."* ## Without MCP Everything is plain HTTP. Any agent that can run `curl` can follow the "start here" workflow at the top of the [API description](https://api.canonopylabs.com/openapi.json), or the [API reference](https://canonopylabs.com/docs/api-reference.md). --- Source: https://canonopylabs.com/docs/godot (Markdown: https://canonopylabs.com/docs/godot.md) # Use your model in Godot Your downloaded model can decide inside your Godot 4 game: from GDScript, offline, every frame if you like. It makes exactly the same decisions as the hosted `/v1/decide` and the [Python runtime](https://canonopylabs.com/docs/running-it-yourself.md), in the same response format, with your rules enforced on every decision. No Python and no ONNX Runtime: the addon reads the model file itself. The addon is a 33 KB download; a Snake model is 175 KB and a Doom model about 800 KB. It works on every platform Godot exports to, web included (Godot 4.3 or later). ## Install Download [canonopy-godot-0.1.1.zip](https://canonopylabs.com/dl/canonopy-godot-0.1.1.zip) and unzip it at the root of your project (it creates `addons/canonopy/`). Then turn it on in **Project > Project Settings > Plugins > Canonopy**, so your model files go into exported games. Want to see it first? [canonopy-godot-example-0.1.1.zip](https://canonopylabs.com/dl/canonopy-godot-example-0.1.1.zip) is a whole Godot project: Snake, played by a model. Open it and press Play. ## Two steps **1. Put your model in the project.** Download it and drop the zip anywhere under your project, for example `res://models/snake@1.zip`. Don't unzip it. ```bash canonopy download snake@1 ``` **2. Ask it, from GDScript.** ```gdscript var model := CanonopyModel.load("res://models/snake@1.zip", {"best_guess": true}) func _physics_process(_delta): var res := model.decide({"state": { "snakeLength": 12, "moves": { "up": {"eats": false, "food_distance": 4, "room": 180, "dead_end": false, "tail_reachable": true}, "right": {"eats": true, "food_distance": 0, "room": 175, "dead_end": false, "tail_reachable": true}, }, }}) var move = res["decision"]["action"] # "right" ``` The state is the same JSON your model was set up with, as a Dictionary. `CanonopyModel.load` returns `null` when it can't load the model, and `CanonopyModel.last_error` says why. ## What comes back The same response as the [hosted API](https://canonopylabs.com/docs/api-reference.md): `answers` per question, and `decision` with your rules applied (`action`, `confidence`, `blocked_by_rules`, `blocked_actions`, `route`). - `route` is `"act"` at or above the model's [safety bar](https://canonopylabs.com/docs/confidence.md), `"review"` below it, and `"ask"` when your rules left no action. With best guess on (below), an answer below the bar says `"act"` too, plus `"best_guess": true`. - `{"questions": {...}}` in the request asks only some questions, or some of a question's options. - On a problem the response is `{"error": {"type", "message"}}`. - `model.decide_many(states)` decides many states at once. ## Best guess (for games) For games, turn on best guess: the model always acts, and tells you how sure it was. ```gdscript var model := CanonopyModel.load("res://models/snake@1.zip", {"best_guess": true}) var res := model.decide({"state": state}) res["decision"]["action"] # always a move your rules allow res["decision"]["confidence"] # how sure it was res["decision"].get("best_guess", false) # true when it was unsure and acted anyway ``` In a game, "unsure" usually means two moves are about as good, not a real doubt. With best guess on, the model plays its most likely move instead of routing it, and marks it: `"route": "act"` plus `"best_guess": true`, with the confidence and probabilities as usual. A confident answer has no `best_guess` field, so you can always tell them apart. Your rules still apply first: a blocked move is never the guess, and when your rules allow no move, `action` is `null` and `route` is `"ask"`. A best guess doesn't call your fallback or the hosted service. It is off by default, because business decisions should keep routing unsure cases to a person. `model.decide(request, Callable(), true)` (or `false`) sets it for one call. The example project has it on and shows "model", "best guess" or "rule" for every move. ## Your own fallback With best guess off, you can answer the moves the model isn't sure of yourself: ```gdscript func my_fallback(state, questions): # only the unsure questions return {"move": safest_move(state)} # an option is enough; null leaves it held for review var res := model.decide({"state": state}, my_fallback) ``` Its answers say `"source": "local_fallback"`, and your rules still apply after them. ## Models that read text A model that reads messages needs the text reader (34 MB and up), too big to ship in a game. Those models decide through the hosted service: add a **CanonopyHosted** node to your scene, then ```gdscript var res = await model.decide_async({"state": {"message": "Where is my order?"}}, $CanonopyHosted) ``` For numbers models, `decide_async` answers locally when the model is sure, and sends only the unsure answers to the service (with best guess on, it acts on them locally instead). If the service can't be reached, they go to your fallback, else they are held for review. > **Never ship your API key inside a game** > Players can read anything inside a released game, a key in a script or a scene included, and anyone with your key can make decisions on your account. While you develop, set `CANONOPY_API_KEY` in your environment before opening Godot: CanonopyHosted reads it, and nothing goes into your project. In a released game, point CanonopyHosted's `base_url` at **your own server** and leave its key empty: your server keeps the key, adds `Authorization: Bearer `, and forwards `POST /v1/decide` (same JSON body) to `https://api.canonopylabs.com/v1/decide`. ## Faster decisions Everything runs in GDScript. For big models deciding every frame, an optional native library runs the model about 50 times faster: unzip [canonopy-godot-native-0.1.1-macos.zip](https://canonopylabs.com/dl/canonopy-godot-native-0.1.1-macos.zip) at the root of your project and restart Godot. Same decisions either way; `model.uses_native()` says which one runs. macOS is built (Apple silicon and Intel); for Windows and Linux, build it from the addon's source (one C++ file, no dependencies). The web uses GDScript. One `decide()` call (reading the state, the model, your rules and the answer), median over 300 recorded game states on a laptop: | Model | GDScript | With the native library | |---|---|---| | Snake (21 numbers, 1 question) | 0.64 ms | 0.14 ms | | Doom (61 fields, 4 questions) | 2.3 ms | 0.50 ms | ## Same decisions as everywhere else Every release is checked against the Python runtime on real models (Snake, Doom, and a numbers-and-categories model with every kind of question and rules on other answers): 2,000 varied requests each, with missing fields, numbers written as text, unknown categories and edge values, with best guess off and on. Chosen answers, routes, best guess flags, rule outcomes and refusals were identical in every case. ## Not yet - Reading text on the device (text models decide through the hosted service). - Saving unsure cases on the device to sync later (the Python runtime's `sync`). - Models downloaded before Sep 26, 2026: download them again. --- Source: https://canonopylabs.com/docs/using-with-llms (Markdown: https://canonopylabs.com/docs/using-with-llms.md) # Using it with your LLM Language models are good at reading between the lines and writing replies. They're not a good place for a rule-bound decision: they can be talked out of your rules, and their confidence isn't calibrated on your cases. Put the decision in your own model, and let the LLM do the talking. ## Support bot Your model picks the action; your LLM writes the reply around it. The bot can't promise a refund your rules don't allow, because the action is decided before the LLM writes a word. ```python res = decide(model="bank-support@latest", state={"message": msg, **account}) d = res["decision"] if d["route"] == "ask": reply = llm(f"Tell the customer a person will pick this up shortly. Message: {msg}") hand_to_person(res["id"]) else: reply = llm( f"You are a bank's support agent. The next step is decided: {d['action']}. " f"Explain it to the customer in two sentences. Don't offer anything else. Message: {msg}" ) ``` ## Guardrails Before an action an LLM proposes is carried out, check it against your model. Act only if no rule blocks it and your model agrees: ```python res = decide(model="refunds@latest", state=case) d = res["decision"] allowed = llm_action not in d["blocked_actions"] and d["action"] == llm_action if not allowed: log("blocked", llm_action, d["blocked_by_rules"]) # rule ids, in words you wrote ``` ## Agent decision layer Give your agent one tool, `decide`, for the choices that must follow your rules. It gets calibrated answers and a route instead of guessing. ```json { "name": "decide", "description": "Decide the next action for a customer case. Follows the company's rules. Returns the action, a confidence and a route: act, review or ask.", "input_schema": { "type": "object", "properties": { "state": { "type": "object", "description": "The case: message and account fields." } }, "required": ["state"] } } ``` Tell the agent: on `act`, go ahead; on `review`, go ahead and flag it; on `ask`, stop and hand over. ## Your LLM as the fuzzy judge Some rules depend on things like tone or urgency. Ask your LLM (or Jev) for that judgment and pass it in as a state field; your model makes the rule-bound decision. Or ask it as a question: a question your model hasn't learned is answered by your [day-one backend](https://canonopylabs.com/docs/day-one.md), and graduates once it has outcomes. ## Your LLM answering the unsure queue The `ask` cases can go to a stronger model instead of a person. Post its answer to `/v1/domains/{domain}/answers`, and the next version learns from it. --- Source: https://canonopylabs.com/docs/patterns (Markdown: https://canonopylabs.com/docs/patterns.md) # Patterns These are Jev's four patterns. They work unchanged on our answers; the difference is that the numbers come from a model that learned your decision, and your rules are already applied. ## Speculative fan-out Ask every question in one request. They're answered together, and you only pay one round trip. ```python res = decide(model="triage@latest", state=ticket, questions={ "team": team_q, "priority": priority_q, "churn_risk": churn_q, }) a = res["answers"] route(a["team"]["choice"], priority=a["priority"]["score"], flag=a["churn_risk"]["noul"] > 0.7) ``` Pricing is flat, so asking more questions costs nothing extra. ## Confidence-gated routing Act on confident answers; send the rest somewhere smarter or to a person. ```python d = res["decision"] if d["route"] == "act": do(d["action"]) elif d["route"] == "review": do(d["action"]); review_later(res["id"]) else: escalate(res["id"]) ``` The safety bar sets itself after every training run, so `route` already says which answers are automatic. See [Confidence](https://canonopylabs.com/docs/confidence.md). ## Composite scoring Combine several answers with weights, in your code: ```python a = res["answers"] risk = 0.5 * a["fraud_likely"]["noul"] + 0.3 * (a["urgency"]["score"] / 2) + 0.2 * a["new_payee"]["noul"] if risk > 0.6: hold(txn) ``` Keep the weights in code where you can test them. If the combination is itself a decision you make every day, make it a domain's decision question instead, and let it learn from outcomes. ## Intent routing Classify first, then hand each case to the right handler: code, a specialist LLM, or a person. ```python topic = res["answers"]["topic"]["choice"] handler = {"refund": refunds_flow, "lost_stolen": block_and_replace, "info": faq_bot}.get(topic, human) handler(ticket) ``` Your rules still apply to the decision question, whatever the handler does next. --- Source: https://canonopylabs.com/docs/cookbooks/doom-zero-data (Markdown: https://canonopylabs.com/docs/cookbooks/doom-zero-data.md) # Doom with zero data, head to head with Jev A community Doom integration asks Jev four questions on every tick: how to move, where to look, whether to fire, and whether to open a door. This cookbook builds your own player from **the request the game already sends Jev**, with no recorded play, through `POST /v1/build`, then plays it against Jev. ## The result The model below was built by `POST /v1/build` from the request in step 1, with no recorded play and no hand-written help. The build took about 4 minutes; the model is 638 KB to download. It played the same 13 games as Jev, on the same map (Freedoom MAP01), with the same state, the same four questions and the same controller. Both sides decided at the same rate: every 500 ms each side was sent the state and its answer was played on the next tick, whatever its answer time. **Against Jev's default style** (its instructions say to preserve health, retreat early and avoid unnecessary fights): | 13 games | Your model, zero data | Jev | |---|---|---| | Kills | **45** | 39 | | Deaths | **2** | 5 | | Map squares explored | **186** | 164 | | Seconds spent walking into walls | 59.7 | **27.3** | | Median answer | **4 ms** | 144 ms | | Games won | **7 of 13** | 6 | **Against Jev's aggressive style** (fight, close distance, never retreat), the same games: | 13 games | Your model, zero data | Jev | |---|---|---| | Kills | **45** | 21 | | Deaths | **2** | 11 | | Map squares explored | **186** | 99 | | Games won | **12 of 13** | 1 | A game is won by staying alive, then by more kills, then by exploring more. > **Read these before quoting the numbers** > - It's one build, and 13 games per style on one map is a small sample; a seed fixes the start, not the whole real-time game. An earlier build, made before two changes to the builder, lost: 17 kills and 13 deaths. > - Jev's side is its recorded games from our run of live Jev at the same decision rate, on the same 13 seeds; ours played the same seeds at the same rate. Jev missed 9 of its 3,015 ticks (an answer not back in time is dropped and the last move repeated); ours missed none. > - Our model was served from the same machine as the game; Jev's answers came over the internet. The 4 ms and 144 ms are each side's own answer time; the game waited 500 ms for both. > - It walked into walls for twice as long as Jev's default style, mostly heading for pickups and doors the state doesn't say it can reach. > - On fresh situations, its four answers agreed with its instructions and rules 95.7% (movement), 100% (view), 99.9% (trigger) and 100% (interaction) of the time. > - For comparison, hand-built (not what a build gives you): a player we built by hand from the same inputs, with no recorded play, got 52 kills and 0 deaths against the default style; a player built from recorded play got 56 kills and 0 deaths. See [Doom with recorded play](https://canonopylabs.com/docs/cookbooks/doom.md). ## Build it ### 1. Take the request the game sends Jev Every tick the game sends Jev its whole structured state and four Choice questions. Their instructions hold the task, the playing style and the integration's constraints, and the options are the integration's own, word for word. Save one real request as `doom-jev-request.json` (the state is shortened here). The build above was made from exactly this request, plus the two optional parts after it: ```json { "model": "jev-latest", "state": { "player": { "health": 100, "armor": 100, "weapon": "pistol", "ammo": { "bullets": 50, "shells": 32, "rockets": 0, "cells": 0 }, "recent_damage": 0, "under_fire": false, "x": -192, "y": -192, "angle": 0, "motion": { "distance": 0, "stuck": false, "total_distance": 0 } }, "visible_enemies": [ { "id": "monster_1", "distance_units": 232, "health": 30, "attacking": true, "threat": "high", "…": "…" } ], "visible_pickups": [ { "id": "pickup_1", "type": "item", "distance_units": 208, "reachable": true, "…": "…" } ], "combat": { "enemy_detected": true, "visible_enemy_count": 3, "nearest_visible_enemy": { "distance": 232, "relative_angle": 0, "aligned": true, "in_fighting_range": true, "…": "…" } }, "exploration": { "visited_cells": 1, "current_cell_visits": 2, "novelty": 0.5 }, "history": { "previous_action": "EXPLORE_WORLD", "current_intent": "survive_level" }, "world": { "entities": [ { "distance": 208, "relative_angle": 960960175, "visible": true, "enemy": false, "pickup": true, "…": "…" } ] } }, "questions": { "movement": { "type": "choice", "instructions": { "task": "Choose navigation for this tick.", "policy": "Preserve health above everything else. Retreat early, use cover, collect health and armor, and avoid unnecessary fights.", "constraints": [ "Choose all four axes independently; they execute simultaneously.", "Approaching and facing do not imply firing. Fire only when a living enemy is visible, aligned, in range, and ammunition is appropriate.", "Use the combat sensor, player condition, exploration memory, entities, and linedefs.", "If stuck, change movement or view; do not idle without a reason." ] }, "criteria": { "HOLD_POSITION": "Do not translate.", "EXPLORE_WORLD": "Navigate toward under-visited space.", "MOVE_TO_ENEMY": "Path toward the nearest living enemy and stop at fighting distance.", "RETREAT_FROM_ENEMY": "Create distance from the nearest enemy.", "COLLECT_NEAREST_PICKUP": "Path toward the nearest useful pickup.", "MOVE_TO_USE": "Path toward a usable door or switch." } }, "view": { "type": "choice", "instructions": { "task": "Choose where to look.", "…": "the same policy and constraints" }, "criteria": { "KEEP_HEADING": "Keep the current heading.", "SCAN": "Turn to inspect the environment.", "FACE_ENEMY": "Center the nearest visible enemy." } }, "trigger": { "type": "choice", "instructions": { "task": "Choose whether to fire.", "…": "the same policy and constraints" }, "criteria": { "HOLD_FIRE": "Do not fire.", "FIRE": "Press the weapon trigger." } }, "interaction": { "type": "choice", "instructions": { "task": "Choose whether to use a nearby line.", "…": "the same policy and constraints" }, "criteria": { "NO_USE": "Do not activate anything.", "USE": "Activate a nearby door or switch." } } } } ``` Send the real thing, not a shortened copy: the whole state as the game sends it, and every question with its full instructions. Two optional parts made the build above better, and cost nothing to add: - **`examples`: more real states.** The build above had 16 more states the game really sent, from games on other seeds than the 13 test games, as a list in `examples`. States only, no answers. - **`fields`: what the numbers mean.** A note for the fields whose units aren't obvious, keyed by path: ```json "fields": { "player.angle": { "meaning": "the direction the player faces, as a Doom binary angle: 0 = east, 1,073,741,824 = north", "unit": "binary angle (2^32 per turn)", "range": [0, 4294967295] }, "combat.nearest_visible_enemy.relative_angle": { "meaning": "direction from the player's facing to the nearest visible monster; positive = to the left; aligned when its size is at most combat.aim_tolerance", "unit": "binary angle (2^32 per turn)" }, "combat.fighting_distance": { "meaning": "the distance within which a monster counts as in fighting range; always 384", "unit": "map units" }, "player.recent_damage": { "meaning": "Doom's damage counter: goes up by the damage just taken; 0 = not hurt recently", "range": [0, 100] }, "history.stuck_probability": { "meaning": "1 when the player has been trying to move but moved under 2 map units for 3 or more decisions in a row, else 0", "range": [0, 1] } } ``` ### 2. Build, with the firing rules in plain words The integration's firing constraint, said as rules, so it can never be broken: These are the rules the build above was given, word for word: ```bash curl https://api.canonopylabs.com/v1/build -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d "$(jq '. + {name: "doom", rules: "Never FIRE when enemy_detected is false. Never FIRE when enemy_aligned is false. Never FIRE when enemy_in_range is false. Never FIRE when weapon is pistol and bullets is 0. Never FIRE when weapon is chaingun and bullets is 0. Never FIRE when weapon is shotgun and shells is 0. Never FIRE when weapon is super_shotgun and shells is 0. FIRE when enemy_detected is true and enemy_aligned is true and enemy_in_range is true. Otherwise HOLD_FIRE."}' doom-jev-request.json)" ``` Or with the CLI: `canonopy build doom-jev-request.json --name doom --rules "Never FIRE when enemy_detected is false. …" --wait`. Check `understood` in the answer: - **`fields`:** every value in the state, each with its `path` and type: the player's health, armour, weapon and ammunition, what's in sight and how lined up and far it is, where it has been, and the nearest things around it (lists of objects are read at their first 8 positions). `not_read` lists what's left out and why, like the ids. - **`decision_question`:** `trigger`, the question the rules are about. Your hard rules apply to it. - **`rules`:** every sentence, each with the field it was matched to (`enemy_aligned` became `combat.nearest_visible_enemy.aligned`). The "Never" sentences are enforced on every decision, whatever the model thinks; the rest your model follows in its answers. `rules_not_understood` is empty. The other three questions follow their instructions (the task, the style and the constraints), which your model learns. ### 3. Wait for `ready` There's nothing to do in between. It's built and checked against your instructions and rules on fresh situations; each of the four questions has to agree at least 95% of the time before it serves. `canonopy build status doom` shows where it is (`--wait` on the build returns when it's ready), and `weakest` shows where each question matches least. See [the quality check](https://canonopylabs.com/docs/start-with-no-data.md#the-quality-check-95-on-fresh-situations). ### 4. See how it decides (optional) Any time once it's ready: about 20 game situations, each with the answer to all four questions and why. ```bash canonopy examples doom canonopy signoff doom --correct ex_04:movement=RETREAT_FROM_ENEMY --retrain # only if one is wrong ``` Look for what you care about in a player: does it retreat when hurt, stop at fighting distance, look around when stuck? A correction with `--retrain` makes a new version that gives that answer. ### 5. Point the game at it Nothing in the integration changes but the endpoint, the key and `model`: ```diff - POST https://api.typesafe.ai/v1/systemone "model": "jev-latest" + POST https://api.canonopylabs.com/v1/systemone "model": "doom@latest" ``` The game keeps sending its whole state and all four questions, and gets an answer to each, in Jev's shape: ```json { "model": "doom@1", "answers": { "movement": { "type": "choice", "choice": "HOLD_POSITION", "confidence": 0.98, "source": "trained", "route": "act", "…": "…" }, "view": { "type": "choice", "choice": "FACE_ENEMY", "confidence": 1.0, "source": "trained", "route": "act", "…": "…" }, "trigger": { "type": "choice", "choice": "FIRE", "confidence": 0.99, "source": "trained", "route": "act", "…": "…" }, "interaction": { "type": "choice", "choice": "NO_USE", "confidence": 1.0, "source": "trained", "route": "act", "…": "…" } }, "decision": { "question": "trigger", "action": "FIRE", "confidence": 0.99, "blocked_by_rules": [], "blocked_actions": [], "route": "act" } } ``` The controller that played Jev's answers plays these unchanged. To run it inside the game instead, next to the engine, download it and use `canonopy-runtime`: see [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md). ## Make it better Watch it play. When it does something you don't want, you have three ways to fix it, and each retrain takes about a minute: - **The rule is wrong or missing:** say it in plain words, like "Never MOVE_TO_ENEMY when enemy_detected is false", with `POST /v1/domains/doom/rules`. It retrains straight away, and the build says what changed. See [Change a rule with a preview](https://canonopylabs.com/docs/improve-your-model.md#change-a-rule-with-a-preview). - **A move was wrong:** mark it wrong in the decision log with the right answers. See [Mark a decision wrong](https://canonopylabs.com/docs/improve-your-model.md#mark-a-decision-wrong). - **You have recorded play:** send it as history and build again under the same name. Moments of play then come first. See [Doom with recorded play](https://canonopylabs.com/docs/cookbooks/doom.md) for what recordings did for this game. --- Source: https://canonopylabs.com/docs/cookbooks/snake-zero-data (Markdown: https://canonopylabs.com/docs/cookbooks/snake-zero-data.md) # Snake with zero data, head to head with Jev Jev's own demo repository has a Snake game that asks Jev for every move. Its request holds the game state, with the board and each legal move's facts (does it eat the food, how far is the food after it, how much room is left, is it a dead end, can the snake still reach its tail), a goal in plain words, and a Choice question over the moves. This cookbook builds your own Snake player from **that request as the demo sends it**, with no recorded moves, through `POST /v1/build`, and runs it inside the game loop. ## The result The model below was built by `POST /v1/build` from the request in step 1, with no recorded moves and no hand-written help, in about 3 minutes. It's 445 KB to download and reads numbers and categories only, so it needs no text reader. The demo's 20 test games on its 16×16 board, each stopped at the step where Jev's game on that seed stopped (about 700 moves), so both sides had the same number of moves: | 20 games | Your model, zero data | Jev, the demo's default `safe` strategy | |---|---|---| | Food eaten, per game | **42.0** ± 4.1 | 40.0 ± 3.3 | | Games that ended in a crash | **0** | 1 | | Games ahead / level / behind Jev | **12 / 2 / 6** | | | Time per move | **0.16 ms**, downloaded, in the game's process | 131 ms, median round trip | Against Jev's `greedy` strategy, Jev ate more: 40.8 against 35.6. It crashed in 13 of those 20 games; our model in none. > **Read these before quoting the numbers** > - 20 games is a small sample. Jev's games were stopped after about 700 moves to cap what the test spent, so the comparison is over the same number of steps. Jev's side is from our recorded run of Jev on the demo's own code; ours played the same engine and the same 20 games. > - This is the best of four automatic builds on crashes. Three of the four ate more than Jev (40.9, 42.0 and 43.4 against 40.0) and one ate less (36.5); the other three crashed in 1, 3 and 5 games. > - 97% of its moves came back with `route: "act"`; the game plays the answer either way, as below. > - Our 0.16 ms is the downloaded model on a laptop CPU, with each move's facts already worked out by the demo; Jev's 131 ms is a round trip over the internet. > - For comparison, hand-built (not what a build gives you): a player we built by hand from the same request ate 47.1 and never crashed. A model built through the API from 3,000 moves recorded from a simple scripted player ate 46.1 and crashed in 3 of 20. See [A Snake bot from recorded moves](https://canonopylabs.com/docs/cookbooks/snake.md). ## Build it ### 1. Take the request the demo sends Jev Save one real request, as the demo sends it, as `snake-jev-request.json` (the board is shortened here): ```json { "model": "jev-latest", "state": { "board": ["........ooooo...", "........o...o...", "…", "...F........H...", "…"], "legend": "H = snake head, o = snake body, T = snake tail, F = food, . = empty. Row 0 is the top row, column 0 is the leftmost column. Leaving the grid hits a wall.", "head": { "row": 8, "col": 12 }, "food": { "row": 8, "col": 3 }, "heading": "down", "snakeLength": 20, "gridSize": { "rows": 16, "cols": 16 }, "foodIsAdjacent": false, "facts": [ { "dir": "right", "turn": "left turn", "target": { "r": 8, "c": 13 }, "eats": false, "foodDistance": 10, "reachable": 236, "freeTotal": 236, "deadEnd": false, "canReachTail": true }, { "dir": "down", "turn": "straight", "target": { "r": 9, "c": 12 }, "eats": false, "foodDistance": 10, "reachable": 236, "freeTotal": 236, "deadEnd": false, "canReachTail": true }, { "dir": "left", "turn": "right turn", "target": { "r": 8, "c": 11 }, "eats": false, "foodDistance": 8, "reachable": 236, "freeTotal": 236, "deadEnd": false, "canReachTail": true } ] }, "questions": { "move": { "type": "choice", "instructions": "You are playing the game Snake. Choose the direction the snake's head moves on the next step. Every listed direction is safe for this single step; the facts next to each direction were computed by code and are exact. A move marked DEAD END almost always loses the game a few steps later. The snake grows by one when it eats the food, and the game ends when the head hits a wall or its own body. Player strategy: Stay alive above all. Prefer the move that keeps the most empty cells reachable, never enter a dead end, and only go for the food when that does not shrink your room. Make no mistakes.", "criteria": { "up": "move up", "down": "move down", "left": "move left", "right": "move right" } } } } ``` Send the whole state, as the demo sends it. The instructions are the demo's rules and its `safe` strategy, word for word. The build above also had a one-line goal in `description` ("Eat as much food as possible without dying.") and 10 more real states in `examples`, from games on other seeds than the 20 test games: states only, no moves. ### 2. Build ```bash curl https://api.canonopylabs.com/v1/build -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d "$(jq '. + {name: "snake", description: "Eat as much food as possible without dying."}' snake-jev-request.json)" ``` Or `canonopy build snake-jev-request.json --name snake --wait`. No rules are needed here: the question's options will be the legal moves, so a move into a wall or the snake is never on offer. `understood.fields` lists what it reads: each board row, the head, the food, the heading, the length, and each move's facts by its direction (`facts_up_deadEnd` is `facts[dir=up].deadEnd`, and so on). `not_read` lists the legend: the same fixed text in every state, so it doesn't help decide. ### 3. Wait for `ready` `canonopy build status snake`, or `--wait` on the build. There's nothing to do in between. It's checked against your instructions on fresh situations, and serves once it agrees at least 95% of the time. See [the quality check](https://canonopylabs.com/docs/start-with-no-data.md#the-quality-check-95-on-fresh-situations). ### 4. See how it decides (optional) ```bash canonopy examples snake canonopy signoff snake --correct ex_07:move=left --retrain # only if one is wrong ``` Each example is a situation with the move your model makes and why. Look for the moves you'd make: take the food when it's safe, never take a dead end. ### 5. Ask for each move Like Jev's demo, list only the legal moves as the question's options; the model picks among them. Send the same state the demo sends Jev: ```bash curl https://api.canonopylabs.com/v1/decide -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d '{ "model": "snake@latest", "state": { "board": ["…"], "head": { "row": 8, "col": 12 }, "food": { "row": 8, "col": 3 }, "heading": "down", "…": "…", "facts": [ { "dir": "right", "…": "…" }, { "dir": "down", "…": "…" }, { "dir": "left", "…": "…" } ] }, "questions": { "move": { "type": "choice", "instructions": "Which way should the snake move next?", "criteria": { "right": "move right", "down": "move down", "left": "move left" } } } }' ``` The move is `answers.move.choice` (with no rules there is no `decision` block: `decision` is for the question your rules are about). ### 6. Run it inside the game loop A move every few milliseconds is where a download pays off: no network, nothing per move. ```bash curl -L -o snake.zip -H "Authorization: Bearer $CANONOPY_API_KEY" \ https://api.canonopylabs.com/v1/models/snake@latest/download ``` ```python from canonopy_runtime import load bot = load("snake.zip") def next_move(state): options = {f["dir"]: f"move {f['dir']}" for f in state["facts"]} # the legal moves only res = bot.decide({"state": state, "questions": {"move": {"type": "choice", "criteria": options}}}) return res["answers"]["move"]["choice"] ``` `pip install https://canonopylabs.com/dl/canonopy_runtime-0.6.3-py3-none-any.whl`. See [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md). ## Make it better - **Recorded moves:** from a scripted player, a designer or your best players, each with the move that was made. Send them as history and build again under the same name. See [A Snake bot from recorded moves](https://canonopylabs.com/docs/cookbooks/snake.md). - **A move you'd never make:** mark it wrong in the decision log with the right move, and retrain. See [Improve your model](https://canonopylabs.com/docs/improve-your-model.md). - **A different style:** change the strategy in the instructions (the demo's `greedy`, say) and build again under a new name, then play the two against each other. --- Source: https://canonopylabs.com/docs/cookbooks/bank-zero-data (Markdown: https://canonopylabs.com/docs/cookbooks/bank-zero-data.md) # Bank support: start in minutes, then improve with your decisions A bank's support inbox: each message is about one of eleven topics and goes to one of nine next steps, under five hard rules. This cookbook builds the desk's model from **the request it would send Jev**, with no past messages, then shows how it pulls ahead as it learns from the desk's decisions, on 3,080 real customer messages. ## The result | 3,080 real messages | Accuracy | |---|---| | **Day one**, built from the Jev request with no data (10 automatic builds) | 85.5% (84.6–86.8%) | | **After about 400 reviewed cases** (the cases it was least sure about, answered) | **91.7%** | | After about 800 reviewed cases | **92.7%** | | **With the bank's history** (about 9,000 past messages) | **about 95%** (94.9–96.0%) | | Jev, few-shot (22 example messages in the prompt) | 86.6% | | Jev, zero-shot (the written policy only) | 83.9% | **Close to Jev on day one; clearly ahead once it learns from your decisions.** A build takes about 5 minutes. On day one it's level with few-shot Jev, within a point or two either way, and ahead of zero-shot Jev. Answer the cases it sends to review and retrain: after about 400 of them it's 5 points ahead of Jev. If you'd rather just report outcomes as they come, about 800 random cases get it to about 90%, and 1,600 to about 92.6%. The model is about 830 KB to download and shares the standard 34 MB English text reader. > **Read these before quoting the numbers** > - The messages are the banking77 test split (PolyAI, CC BY 4.0). It's public, so either model may have seen it before. > - Accuracy is against a written routing policy, the same for both sides. The account details attached to each message were assigned at random for the test; only the text is real. > - Day one: the 9 builds were made by `POST /v1/build` on Sep 30 with slightly different settings while we tuned it. Two more builds with a bug since fixed, and two experiments with other settings, scored 81.6–86.75% and aren't in the range. > - The reviewed cases came from the bank's real past messages (never the test ones), answered with their true answers, as a person reviewing them would; each step is the mean of 2 runs. That test started from a no-data model made from hand-written example messages (88.6% before any cases), a stronger start than today's automatic builds. Answers that come mostly from a backup model, not a person, will likely help less. > - Below about 200 cases, don't expect a visible change: steps of 25 to 100 cases move it less than the ±1 point between two trainings. > - Rules on your fields (amounts, verification, the fraud flag) are never broken. Rules on what a message is about, like "block a lost card first", are enforced whenever there's a real chance they apply, and unsure cases go to review; on day one some messages about lost cards are read as something else, so a few of those cases can be missed. Fewer are missed as it learns; with the bank's full history and those rules, the test had none. > - Long messages are different: on real consumer complaints (CFPB narratives, a few hundred words each), a model built with no data scored about 40% against Jev's 63%. With the complaints' history it scored 77.7% against Jev's 65.0% few-shot. For long texts, start from your past cases. ## Build it ### 1. Take the request you'd send Jev One real example of the state (the message and the account fields), and the two questions with a one-line description of each option. Save it as `bank-jev-request.json`: ```json { "model": "jev-latest", "state": { "message": "I think someone has my card, there are payments I did not make", "tier": "plus", "account_age_days": 812, "amount": 86.5, "prior_disputes": 0, "fraud_flag": false, "card_status": "active", "identity_verified": true, "contacts_7d": 1, "new_device_login_24h": false }, "questions": { "topic": { "type": "choice", "instructions": "What is the customer writing about?", "criteria": { "lost_stolen": "lost, stolen or compromised card (or a lost or stolen phone with the app on it)", "unrecognised": "a payment, cash withdrawal or direct debit they don't recognise, a double charge, or the ATM gave the wrong amount of cash", "refund": "they want a refund from a merchant, or a refund hasn't shown up", "card_fault": "their card doesn't work: declined, contactless or virtual card failing, PIN blocked, or the ATM swallowed it", "card_delivery": "getting, ordering, activating, linking or replacing a card, card arrival, changing the PIN", "payments": "a bank transfer or payment is pending, failed, declined, reverted, cancelled or hasn't arrived, or how to send or receive money", "topup": "topping up the account: a top-up is pending, failed or reverted, top-up methods and limits, cash or cheque deposits not showing", "fees": "a fee or charge they were billed (card payment, transfer, top-up, cash withdrawal, or an extra charge on the statement)", "fx": "exchange rates, currency exchange, a wrong exchange rate, or which currencies are supported", "identity": "verifying their identity or source of funds, a forgotten passcode, or editing personal details", "info": "general product questions: where the card is accepted, Visa or Mastercard, supported countries, ATMs, Apple Pay or Google Pay, card limits, age limits, closing the account" } }, "action": { "type": "choice", "instructions": "What should the support desk do?", "criteria": { "block-card": "block the customer's card straight away so nobody can use it", "refund": "refund the money to the customer", "open-dispute": "open a chargeback dispute on the transaction", "escalate-fraud": "hand the case to the fraud team", "route-payments": "send the case to the payments and transfers team", "route-cards": "send the case to the cards team", "request-verification": "ask the customer to verify their identity first", "answer-faq": "answer the question directly from the help centre", "escalate-senior": "hand the case to a senior human agent" } } } } ``` The option descriptions matter most here: they're how your model knows what each topic covers. ### 2. Build, with the desk's policy in plain words Save the desk's policy, one rule per sentence, as `policy.txt`: ```text Never refund when amount is over 250. Never refund when identity_verified is false. When fraud_flag is true, only escalate-fraud or block-card. Always block-card when topic is lost_stolen and card_status is active or frozen. Never answer-faq when topic is unrecognised. escalate-fraud when fraud_flag is true. route-cards when topic is lost_stolen. request-verification when topic is identity. request-verification when identity_verified is false and topic is one of refund, unrecognised, payments, topup. escalate-senior when topic is unrecognised and prior_disputes is at least 3. open-dispute when topic is unrecognised. answer-faq when topic is fees and tier is not one of premium, metal. escalate-senior when topic is refund or fees and amount is over 250. escalate-senior when topic is refund or fees and account_age_days is under 30. refund when topic is refund or fees. escalate-senior when topic is one of card_fault, card_delivery, payments, topup and contacts_7d is at least 3. route-cards when topic is card_fault or card_delivery. route-payments when topic is payments or topup. Otherwise answer-faq. ``` ```bash curl https://api.canonopylabs.com/v1/build -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d "$(jq --rawfile rules policy.txt '. + {name: "bank-support", rules: $rules}' bank-jev-request.json)" ``` Or `canonopy build bank-jev-request.json --name bank-support --rules "$(cat policy.txt)" --wait`. In `understood`: - **`fields`:** `message` (text) and the nine account fields: `tier` and `card_status` as categories, `fraud_flag`, `identity_verified` and `new_device_login_24h` as yes/no, the rest as numbers. - **`decision_question`:** `action`. - **`rules`:** the first five sentences, which block answers, are enforced on every decision. The first three read account fields, so they hold whatever the model thinks. The two about `topic` depend on what the message is about: they're enforced whenever there's a real chance they apply, and unsure cases go to review. See [Rules on what a message is about](https://canonopylabs.com/docs/decisions-and-rules.md#rules-on-what-a-message-is-about). - The routing sentences ("route-cards when topic is card_fault or card_delivery") aren't hard rules; your model follows them in its answers. ### 3. Wait for `ready`, then switch over There's nothing to do in between: `canonopy build status bank-support` shows where it is. Once it's `ready`, change the base URL and `model`, and keep sending the same JSON: ```diff - POST https://api.typesafe.ai/v1/systemone "model": "jev-latest" + POST https://api.canonopylabs.com/v1/systemone "model": "bank-support@latest" ``` Each answer says where it came from and whether to act on it; `decision.blocked_by_rules` says which rules stopped which actions. A message it's unsure about comes back with `route: "review"`: send those to a person, and their answers teach the next retrain the most. ### 4. See how it decides (optional) Any time once it's ready: about 20 messages with their account details, each with the topic and action your model gives, and why. ```bash canonopy examples bank-support canonopy signoff bank-support --correct ex_05:topic=unrecognised --retrain # only if one is wrong ``` ## Make it better with your decisions This is where it pulls ahead. Two ways, and you can do both. **Answer the cases it was unsure about.** Every message it isn't sure of comes back with `route: "review"` and waits in the unsure queue. Answer them in the console (Improve) or the API, then retrain; `GET /v1/domains/bank-support/advice` says which cases help most. Those are the reviewed cases above: about 400 of them took it to 91.7%. See [Improve your model](https://canonopylabs.com/docs/improve-your-model.md). **Send your history.** The desk's past messages, each with the topic or action it got. Build again under the same name: ```bash canonopy build bank-jev-request.json --name bank-support --rules "$(cat policy.txt)" --history past_messages.csv --wait ``` - Your past messages become the main examples, and your model is measured on some it never saw: the report says how it does **on your own cases**. - **Rules found in your history** are proposed for you to confirm, each with how many past cases it covers. Only the ones you confirm are used. See [Rules found in your history](https://canonopylabs.com/docs/rules-found.md). - Past messages that contradict your rules are set aside for you to review, never learned silently. With the bank's history, the same kind of model scored about 95% (94.9–96.0%). The numbers, and building from history alone: [Bank support with history](https://canonopylabs.com/docs/cookbooks/bank-support.md). --- Source: https://canonopylabs.com/docs/cookbooks/bank-support (Markdown: https://canonopylabs.com/docs/cookbooks/bank-support.md) # Bank support with history, head to head with Jev A bank's support inbox: each message goes to one of nine next steps, under five hard rules. We ran the same 3,080 real customer messages through hosted Jev and through a model learned from the bank's cases. This cookbook shows the numbers, then how to build a model like it from your history. No history yet? Build the same decision from its Jev request in minutes: close to Jev on day one (85.5% against 86.6%), and 91.7% after about 400 reviewed cases. See [Bank support, from day one](https://canonopylabs.com/docs/cookbooks/bank-zero-data.md). Your history takes it further. ## The result | | Canonopy | Jev, few-shot | Jev, zero-shot | |---|---|---|---| | Accuracy | **96.0%** ± 0.4 | 86.6% | 83.9% | | Calibration error | **0.006** | 0.044 | 0.063 | | Time per decision (p50) | **38.9 ms** on a laptop CPU; 0.49 ms once the text is read | 177 ms over HTTPS | 216 ms over HTTPS | | Cost per million decisions | **$0.78** compute | $158.61 at list price | $55.58 at list price | **How much history this took.** Learned from history alone, with no zero-data start: from 1,000 cases it scored 78.3%, below Jev; from 3,000, 88.2%; from 9,000, 94.4%; from 27,000, 96.0%. A case is one message with its account details; the 27,000 are about 9,000 real messages, each with three account states. Starting from your Jev request instead, you don't wait for that history: the model built with no data is close to Jev from day one, and every case you review makes it better from there (91.7% after about 400). See [Bank support, from day one](https://canonopylabs.com/docs/cookbooks/bank-zero-data.md). > **Read these before quoting the numbers** > - The messages are the banking77 test split (PolyAI, CC BY 4.0). It's public, so either model may have seen it before. > - Accuracy is against a written routing policy. Some of Jev's misses were defensible routing calls; counting those as right, the lead is still about 5–7 points. > - The account details attached to each message were assigned at random for the test. Only the text is real; real account data is still untested. > - Our times and costs were measured on a laptop CPU; Jev's are HTTPS wall time and list price. ## Build it This builds a model like it from your own history: past messages with their account details and what the desk did. ### 1. Describe the decision ```bash curl https://api.canonopylabs.com/v1/domains -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d '{"name": "bank-support"}' curl https://api.canonopylabs.com/v1/domains/bank-support/agent -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d '{"message": "We route messages for a bank support desk. Actions: block-card, refund, open-dispute, escalate-fraud, route-payments, route-cards, request-verification, answer-faq, escalate-senior. Fields: message (text), tier (category: basic, plus, premium, metal), amount (number), fraud_flag (yes/no), card_status (category: active, frozen, blocked), identity_verified (yes/no), prior_disputes (number), contacts_7d (number), account_age_days (number). Never refund when amount is over 250; escalate-senior instead. Never refund when identity_verified is false; request-verification instead. When fraud_flag is true, only escalate-fraud or block-card. Otherwise answer-faq."}' ``` The reply: `Set up 9 actions, 9 state fields and 3 rules that can't be broken. Next, train it (looking at the example cases is optional, any time).` Those three rules read account fields, so they're enforced on every decision. The benchmark's other two rules (block a lost card before anything else; never answer an unrecognised payment from the FAQ) depend on what the message is about. Give the domain a `topic` question and say them as rules ("Always block-card when topic is lost_stolen and card_status is active or frozen. Never answer-faq when topic is unrecognised.") and they're enforced whenever there's a real chance they apply, with unsure cases sent to review. See [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md#rules-on-what-a-message-is-about). ### 2. Upload your history One row per past message: its account fields, and the action the desk took, in a column named `action`. ``` message,tier,amount,fraud_flag,card_status,identity_verified,prior_disputes,contacts_7d,account_age_days,action "I lost my card on the bus",plus,40,0,active,1,0,1,812,block-card "I was charged twice for the same coffee",premium,312.4,0,active,1,0,0,1460,escalate-senior ``` ```bash curl https://api.canonopylabs.com/v1/domains/bank-support/data -H "Authorization: Bearer $CANONOPY_API_KEY" -F file=@tickets.csv ``` ### 3. Train ```bash curl -X POST https://api.canonopylabs.com/v1/domains/bank-support/train -H "Authorization: Bearer $CANONOPY_API_KEY" ``` Optional, any time: look at how it decides on 20 example cases. Since you uploaded first, they're your own messages. Open **Review examples** in the console, or use `GET /v1/domains/bank-support/examples`; to correct any, `POST /v1/domains/bank-support/signoff` with `"retrain": true` (see [Training and reports](https://canonopylabs.com/docs/training-and-reports.md#review-the-examples-optional)). ### 4. Call both, side by side ```python import os, requests ACTIONS = { "block-card": "block the card straight away", "refund": "refund the money", "open-dispute": "open a dispute with the merchant", "escalate-fraud": "hand to the fraud team", "route-payments": "send to the payments team", "route-cards": "send to the cards team", "request-verification": "ask the customer to verify their identity", "answer-faq": "answer from the help centre", "escalate-senior": "hand to a senior agent", } state = {"message": "I was charged twice for the same coffee", "tier": "premium", "amount": 312.4, "fraud_flag": False, "identity_verified": True} questions = {"action": {"type": "choice", "instructions": "What should the support desk do?", "criteria": ACTIONS}} ours = requests.post("https://api.canonopylabs.com/v1/systemone", headers={"Authorization": f"Bearer {os.environ['CANONOPY_API_KEY']}"}, json={"model": "bank-support@latest", "state": state, "questions": questions}).json() jev = requests.post("https://api.typesafe.ai/v1/systemone", headers={"Authorization": f"Bearer {os.environ['TYPESAFE_API_KEY']}"}, json={"model": "jev-latest", "state": state, "questions": questions}).json() print(ours["answers"]["action"]["choice"], ours["decision"]["blocked_by_rules"]) print(jev["answers"]["action"]["choice"]) ``` Same request, same answer shape. Ours adds `source` and `route` on each answer and the `decision` block. Here the amount is over 250, so the first rule fires: `blocked_by_rules` is `["never-refund-when-amount-is-over-250-escalate-senior-instead"]`, and `refund` can't be the answer whatever the model thinks. The rule's id comes from its wording; `GET /v1/domains/bank-support` lists them. ### 5. Read the report Each training report shows accuracy on your held-back messages, the weakest actions, the pairs it confuses most with real messages, and a ranked list of what would help. In our benchmark the weakest topic was lost or stolen cards (89%), and most mistakes were between `answer-faq`, `route-cards` and `route-payments`, on questions a reasonable reader could route either way. --- Source: https://canonopylabs.com/docs/cookbooks/doom (Markdown: https://canonopylabs.com/docs/cookbooks/doom.md) # Doom with recorded play, head to head with Jev A community Doom integration asks Jev four questions on every tick: how to move, where to look, whether to fire, and whether to open a door. We pointed the same integration at a Canonopy decision model by changing the URL, the key and `model`, and built that model through this API with an ordinary API key, from recorded play. This cookbook shows the results, then every step of the build. No recorded play? Build it from the game's Jev request alone: built automatically with zero data, it still beat Jev at the same decision rate, 45 kills to 39 and 2 deaths to 5. See [Doom with zero data](https://canonopylabs.com/docs/cookbooks/doom-zero-data.md). Recorded play, as below, took it further. ## The result The same 13 live games for every side, on the same map (Freedoom MAP01), with the same state, the same four questions and the same controller: | | Canonopy `doom-4` | Jev | |---|---|---| | Kills (13 games) | **56** | 34 | | Deaths | **0** | 6 | | Map squares explored, all games | **282** | 114 | | Seconds spent walking into walls | **8.4** | 29.2 | | Median answer | **10 ms** | 125 ms | | Games won | **12 of 13** | 1 | `doom-4` learned from 18,237 recorded moments of play. It took four rounds to get here, each fixing one habit we saw by watching the games: chasing monsters it couldn't reach, walking into walls, and stopping to stare after a fight. Every round was better recordings and a retrain of about 20 seconds. > **Read these before quoting the numbers** > - Our model was served from the same machine as the game; Jev's answers came over the internet. So the times aren't like for like; running the model next to your game is an option Jev doesn't have. > - Thirteen games is a small sample, and Jev's answers vary from run to run: across all our runs on these games it had between 26 and 34 kills and 6 to 8 deaths. > - Kills stop near 4 on this map: the other monsters sit behind doors that neither side got through in the time limits. > - A wall bump is a moment when a side holds forward but hardly moves, with no monster close. Most of Jev's came from walking toward pickups it couldn't reach. > - The side-by-side viewer is a private demo. ## How it works The game sends the same request it sent to Jev: its structured state and four Choice questions. Nothing in the integration changed but the endpoint, the key and `model`: ```diff - POST https://api.typesafe.ai/v1/systemone "model": "jev-latest" + POST https://api.canonopylabs.com/v1/systemone "model": "doom-2@latest" ``` Behind that URL is a normal decision model: - **four questions**, exactly as the game asks them: `movement` (6 options), `view` (3), `trigger` (2) and `interaction` (2); - **`trigger` is the decision question**, so the firing rules are hard rules: it can't fire at nothing, off target, out of range or with no ammunition; - **the state fields** it reads from the game's state, by path: health, ammunition, what's in sight, how lined up and how far the nearest monster is, where it has been, and the nearest things around it. ## Build it ### 1. Create the domain Save this as `doom.json`. The options are the integration's own, word for word. (Here the domain is `doom`; ours in the results is `doom-2`, the second model: see [the lesson](#the-lesson-it-learns-what-your-cases-show).) ```json { "name": "doom", "description": "Doom player controls for each tick", "questions": { "movement": { "type": "choice", "instructions": "Choose navigation for this tick.", "criteria": { "HOLD_POSITION": "Do not translate.", "EXPLORE_WORLD": "Navigate toward under-visited space.", "MOVE_TO_ENEMY": "Path toward the nearest living enemy and stop at fighting distance.", "RETREAT_FROM_ENEMY": "Create distance from the nearest enemy.", "COLLECT_NEAREST_PICKUP": "Path toward the nearest useful pickup.", "MOVE_TO_USE": "Path toward a usable door or switch." } }, "view": { "type": "choice", "instructions": "Choose where to look.", "criteria": { "KEEP_HEADING": "Keep the current heading.", "SCAN": "Turn to inspect the environment.", "FACE_ENEMY": "Center the nearest visible enemy." } }, "trigger": { "type": "choice", "instructions": "Choose whether to fire.", "criteria": { "HOLD_FIRE": "Do not fire.", "FIRE": "Press the weapon trigger." } }, "interaction": { "type": "choice", "instructions": "Choose whether to use a nearby line.", "criteria": { "NO_USE": "Do not activate anything.", "USE": "Activate a nearby door or switch." } } }, "decision_question": "trigger", "state_schema": { "fields": [ { "name": "health", "type": "number", "path": "player.health" }, { "name": "armor", "type": "number", "path": "player.armor" }, { "name": "weapon", "type": "category", "path": "player.weapon", "values": ["fist", "pistol", "shotgun", "chaingun", "rocket_launcher", "plasma", "bfg", "chainsaw", "super_shotgun"] }, { "name": "bullets", "type": "number", "path": "player.ammo.bullets" }, { "name": "shells", "type": "number", "path": "player.ammo.shells" }, { "name": "under_fire", "type": "number", "path": "player.under_fire", "description": "yes/no" }, { "name": "stuck", "type": "number", "path": "player.motion.stuck", "description": "yes/no" }, { "name": "enemy_detected", "type": "number", "path": "combat.enemy_detected", "description": "yes/no" }, { "name": "enemies_in_sight", "type": "number", "path": "combat.visible_enemy_count" }, { "name": "enemy_distance", "type": "number", "path": "combat.nearest_visible_enemy.distance" }, { "name": "enemy_angle", "type": "number", "path": "combat.nearest_visible_enemy.relative_angle" }, { "name": "enemy_aligned", "type": "number", "path": "combat.nearest_visible_enemy.aligned", "description": "yes/no" }, { "name": "enemy_in_range", "type": "number", "path": "combat.nearest_visible_enemy.in_fighting_range", "description": "yes/no" }, { "name": "cell_visits", "type": "number", "path": "exploration.current_cell_visits" }, { "name": "novelty", "type": "number", "path": "exploration.novelty" }, { "name": "previous_action", "type": "category", "path": "history.previous_action", "values": ["EXPLORE_WORLD", "MOVE_TO_ENEMY", "RETREAT_FROM_ENEMY", "FACE_ENEMY", "FIRE", "COLLECT_NEAREST_PICKUP", "USE_NEAREST_LINE", "IDLE"] }, { "name": "pickup_1_distance", "type": "number", "path": "visible_pickups[0].distance_units" }, { "name": "thing_1_distance", "type": "number", "path": "world.entities[0].distance" }, { "name": "thing_1_angle", "type": "number", "path": "world.entities[0].relative_angle" }, { "name": "thing_1_is_enemy", "type": "number", "path": "world.entities[0].enemy", "description": "yes/no" }, { "name": "thing_1_visible", "type": "number", "path": "world.entities[0].visible", "description": "yes/no" } ] } } ``` ```bash curl https://api.canonopylabs.com/v1/domains -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d @doom.json ``` That's a working start. `doom-2` read 74 fields: these, the rest of the player's state (rockets, cells, recent damage, how far it moved), and the same five facts for each of the eight nearest things (`world.entities[0]` to `world.entities[7]`). Everything else in the state, like the long text description and the raw map lines, isn't read. ### 2. Say the firing rules in plain words ```bash curl https://api.canonopylabs.com/v1/domains/doom/agent -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d '{"message": "Never FIRE when enemy_detected is false. Never FIRE when enemy_aligned is false. Never FIRE when enemy_in_range is false. Never FIRE when weapon is pistol and bullets is 0. Never FIRE when weapon is chaingun and bullets is 0. Never FIRE when weapon is shotgun and shells is 0. Never FIRE when weapon is super_shotgun and shells is 0. FIRE when enemy_detected is true and enemy_aligned is true and enemy_in_range is true. Otherwise HOLD_FIRE."}' ``` These are the integration's own constraint ("fire only when a living enemy is visible, aligned, in range, and ammunition is appropriate"), written as rules. The reply: ``` Set up 2 actions, 21 state fields and 7 rules that can't be broken. Next, train it (looking at the example cases is optional, any time). ``` > **Hard rules apply to the decision question** > We also asked for "Never MOVE_TO_ENEMY when enemy_detected is false". The agent refused it: `'MOVE_TO_ENEMY' isn't one of your actions (HOLD_FIRE, FIRE)`. Hard rules apply to the domain's decision question, here `trigger`. The other three questions are shaped by your cases instead (step 3). ### 3. Upload recorded play Your cases are moments from recorded play, each with the right answer to all four questions. They can come from a scripted bot, a designer playing, or your strongest players. One JSON line per moment: the state the game sent, and the answers. ```json {"state": {"player": {"health": 100, "...": "..."}, "combat": {"...": "..."}, "...": "..."}, "answers": {"movement": "MOVE_TO_ENEMY", "view": "FACE_ENEMY", "trigger": "FIRE", "interaction": "NO_USE"}} ``` ```bash curl https://api.canonopylabs.com/v1/domains/doom/data -H "Authorization: Bearer $CANONOPY_API_KEY" \ -F file=@recorded_play.jsonl ``` ```json { "accepted": 6278, "rejected": 0, "total_cases": 6278, "per_question": { "movement": 6278, "view": 6278, "trigger": 6278, "interaction": 6278 }, "problems": [] } ``` ### 4. Train ```bash curl -X POST https://api.canonopylabs.com/v1/domains/doom/train -H "Authorization: Bearer $CANONOPY_API_KEY" ``` `doom-2` took 22 seconds. On its held-back moments: ``` movement trained 97.2% on 675 held-back view trained 98.4% on 675 held-back trigger trained 99.9% on 675 held-back interaction trained 100.0% on 675 held-back rules: 0 violations ``` The report's to-do list was all boundaries, like `SCAN` vs `KEEP_HEADING` and `EXPLORE_WORLD` vs `COLLECT_NEAREST_PICKUP`, and "more cases" items. See [Training and reports](https://canonopylabs.com/docs/training-and-reports.md#the-report). Optional, any time: see how it decides on 20 examples. Since you uploaded first, most of them are real moments from your recordings. ```bash curl https://api.canonopylabs.com/v1/domains/doom/examples -H "Authorization: Bearer $CANONOPY_API_KEY" ``` If one is wrong, correct it with `POST /v1/domains/doom/signoff` and `"retrain": true`, as in [Training and reports](https://canonopylabs.com/docs/training-and-reports.md#review-the-examples-optional). For `doom-2` all 20 came from the recordings and followed the firing rule, so there was nothing to correct. ### 5. Decide The game sends its whole state and all four questions, with their full instructions, in Jev's format. A short request you can try, asking only `trigger`, as `request.json`: ```json { "model": "doom@latest", "state": { "player": { "health": 100, "armor": 100, "weapon": "pistol", "ammo": { "bullets": 50, "shells": 32 }, "under_fire": false, "motion": { "stuck": false } }, "combat": { "enemy_detected": true, "visible_enemy_count": 3, "nearest_visible_enemy": { "distance": 232, "relative_angle": 0, "aligned": true, "in_fighting_range": true } }, "exploration": { "current_cell_visits": 2, "novelty": 0.5 }, "history": { "previous_action": "EXPLORE_WORLD" } }, "questions": { "trigger": { "type": "choice", "instructions": "Choose whether to fire.", "criteria": { "HOLD_FIRE": "Do not fire.", "FIRE": "Press the weapon trigger." } } } } ``` Send the whole state in play: a field left out counts as missing, and the model is less sure without it. ```bash curl https://api.canonopylabs.com/v1/systemone -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d @request.json ``` The game's full request gets an answer to each question: ```json { "model": "doom@1", "answers": { "movement": { "type": "choice", "choice": "HOLD_POSITION", "confidence": 1.0, "source": "trained", "route": "act", "...": "..." }, "view": { "type": "choice", "choice": "FACE_ENEMY", "confidence": 1.0, "source": "trained", "route": "act", "...": "..." }, "trigger": { "type": "choice", "choice": "FIRE", "confidence": 1.0, "source": "trained", "route": "act", "...": "..." }, "interaction": { "type": "choice", "choice": "NO_USE", "confidence": 1.0, "source": "trained", "route": "act", "...": "..." } }, "decision": { "question": "trigger", "action": "FIRE", "confidence": 1.0, "blocked_by_rules": [], "blocked_actions": [], "route": "act" } } ``` The questions a request sends only need the same options as the domain's; their instructions can be as long as the game likes. Each answer keeps Jev's fields, so the controller that played Jev's answers plays these unchanged. To run it inside the game instead, download it and use `canonopy-runtime`: see [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md). ### With the CLI The same steps with the [`canonopy` command](https://canonopylabs.com/docs/cli-reference.md): ```bash canonopy describe doom - < rules.txt # the sentences from step 2 canonopy upload doom recorded_play.jsonl canonopy train doom canonopy report doom canonopy examples doom # optional: how it decides on 20 examples canonopy decide doom@latest --state @state.json --questions @questions.json # the game's state and questions ``` ## The lesson: it learns what your cases show The first model, `doom@1`, never died, but it explored less than Jev: 28 kills and 8.4 squares over the same nine games. After the opening fight it kept answering `MOVE_TO_ENEMY` toward a monster hidden behind a wall, and walked into the wall. That habit was in its recordings. In more than half of the recorded moments with no monster in sight, the recorded answer was to head for the nearest monster anyway. The model learned exactly what it was shown. The fix was better recordings, not a different API. The new recordings showed: - with nothing in sight, never head for a hidden monster: pick up what's close, otherwise explore; - lingering in one spot counts as stuck: look around and move on; - while exploring, press use, so doors open; - in a fight, back off earlier when outnumbered. A retrain takes about a minute. Replace the recordings and train the same model: ```bash curl "https://api.canonopylabs.com/v1/domains/doom/data?mode=replace" -H "Authorization: Bearer $CANONOPY_API_KEY" \ -F file=@better_play.jsonl curl -X POST https://api.canonopylabs.com/v1/domains/doom/train -H "Authorization: Bearer $CANONOPY_API_KEY" ``` Or `canonopy upload doom better_play.jsonl --replace`, then `canonopy train doom`. The old recordings are deleted, and `doom@2` learns from the new ones only. The game keeps calling `doom@latest`, `doom@1` stays available to roll back to, and the report's `cases` says how many earlier cases were removed. The comparison with `doom@1` uses the new recordings' held-back moments, so recordings that fix a bad habit can win it. To drop only some recordings, list the uploads (`canonopy uploads doom`) and delete one; see [Replace or remove cases](https://canonopylabs.com/docs/training-and-reports.md#replace-or-remove-cases). We built our second round before cases could be replaced, as a new domain, `doom-2`, with the same questions and rules. It kept `doom@1`'s zero deaths, explored about 2.5 times as much, and beat Jev on kills, deaths and exploration. A third round, `doom-3`, taught it to turn before walls: wall bumps fell from 16 seconds to 3 over 13 games, with no deaths. But it stopped to stare after fights, so a fourth round, `doom-4`, recorded finishing the fight and then sweeping the level, turning while walking: it explored the most of any version, never died, and won 12 of 13 against Jev. > **A bad habit in the recordings becomes a bad habit in play** > Watch the games, find the moment it goes wrong, and look at what your recordings say to do there. Record better cases for that moment and train again. The report's [to-do list](https://canonopylabs.com/docs/training-and-reports.md#3-what-would-make-it-better) points at the boundaries where more cases help most. --- Source: https://canonopylabs.com/docs/cookbooks/snake (Markdown: https://canonopylabs.com/docs/cookbooks/snake.md) # A Snake bot from recorded moves Jev's own demo repository has a Snake game that asks Jev for every move. Its request holds the board and a Choice question over the legal moves, with facts about each one: does it eat the food, how far is the food after it, how much room is left, is it a dead end, can the snake still follow its tail out. This cookbook plays the same game with your own decision model, learned from recorded moves: the game state in, a Choice question over the moves out. It's a normal decision model, so you can call it like Jev or download it and run it inside the game loop. No recorded moves? Build it from the game's Jev request alone: built automatically with zero data, it ate 42.0 to Jev's 40.0 at equal steps and never died. See [Snake with zero data](https://canonopylabs.com/docs/cookbooks/snake-zero-data.md). ## The results **Built through the API exactly as below**, from 3,000 moves recorded from a simple scripted player, it trained in a few seconds. On the demo's 20 test games, each stopped at the step where Jev's game on that seed stopped: | | Your model | Jev | |---|---|---| | Food eaten, per game | **46.1** | 40.0 | | Games that ended in a crash | 3 of 20 | 1 of 20 | | Time per move | **0.08 ms**, downloaded, in the game's process | 131 ms, median round trip | It eats faster than Jev but crashes more often. It chose the recorded player's move 99% of the time on held-back moves, and the player itself crashed in none of these games; the few moves it gets wrong come in tight spots, where one wrong move ends the game. The fix is more recorded moves from those spots: play the model, find where it goes wrong, record the right move there, and train again (see [the lesson in the Doom cookbook](https://canonopylabs.com/docs/cookbooks/doom.md#the-lesson-it-learns-what-your-cases-show)). In an earlier test of ours, a 6.5 KB model out-ate Jev at equal steps (49–55 food against 40), deciding in well under a millisecond on a laptop against Jev's 131 ms round trip. That model was trained offline in our own test, not through this API. > **Read these before quoting the numbers** > - 20 games on the demo's 16×16 board. Jev played the demo's default `safe` strategy; its games were stopped after about 700 moves to cap what the test spent, so they're compared over the same number of steps. > - A second build from the same moves ate 44.9 and crashed in 5 of 20: each build holds back a different tenth of your cases, so results vary a little. > - Jev's numbers are from our earlier run of Jev on the demo's own code. Ours played the same engine and the same 20 games. > - Our 0.08 ms is the downloaded model on a laptop CPU, facts already computed; Jev's 131 ms is a round trip over the internet. ## Build it ### 1. Send the state, plus each move's facts Send the state you already send Jev. Add each legal move's facts under `moves`, as numbers and yes/no values; a move that isn't legal has no entry. Fields the model doesn't read, like the board, are ignored. ```json { "heading": "right", "snakeLength": 3, "head": { "row": 8, "col": 8 }, "food": { "row": 0, "col": 13 }, "moves": { "up": { "eats": false, "food_distance": 12, "room": 253, "dead_end": false, "tail_reachable": true }, "right": { "eats": false, "food_distance": 12, "room": 253, "dead_end": false, "tail_reachable": true }, "down": { "eats": false, "food_distance": 14, "room": 253, "dead_end": false, "tail_reachable": true } } } ``` ### 2. Create the domain One Choice question, `move`, which is also the decision question, so your rules apply to it. The fields read each move's facts by path. The four rules block a move that has no facts, so as long as you send facts only for legal moves, the snake never moves into a wall or itself, whatever the model thinks. Save this as `snake.json`: ```json { "name": "snake", "description": "Which way the snake moves next", "questions": { "move": { "type": "choice", "instructions": "Which way should the snake move next?", "criteria": { "up": "move up", "down": "move down", "left": "move left", "right": "move right" } } }, "decision_question": "move", "rules": [ { "id": "up-only-when-safe", "description": "Never move up into a wall or the snake", "when": [{ "field": "up_room", "op": "is_missing" }], "block": ["up"] }, { "id": "down-only-when-safe", "description": "Never move down into a wall or the snake", "when": [{ "field": "down_room", "op": "is_missing" }], "block": ["down"] }, { "id": "left-only-when-safe", "description": "Never move left into a wall or the snake", "when": [{ "field": "left_room", "op": "is_missing" }], "block": ["left"] }, { "id": "right-only-when-safe", "description": "Never move right into a wall or the snake", "when": [{ "field": "right_room", "op": "is_missing" }], "block": ["right"] } ], "state_schema": { "fields": [ { "name": "length", "type": "number", "path": "snakeLength" }, { "name": "up_eats", "type": "number", "path": "moves.up.eats", "description": "yes/no" }, { "name": "up_food_distance", "type": "number", "path": "moves.up.food_distance" }, { "name": "up_room", "type": "number", "path": "moves.up.room" }, { "name": "up_dead_end", "type": "number", "path": "moves.up.dead_end", "description": "yes/no" }, { "name": "up_tail_reachable", "type": "number", "path": "moves.up.tail_reachable", "description": "yes/no" }, { "name": "down_eats", "type": "number", "path": "moves.down.eats", "description": "yes/no" }, { "name": "down_food_distance", "type": "number", "path": "moves.down.food_distance" }, { "name": "down_room", "type": "number", "path": "moves.down.room" }, { "name": "down_dead_end", "type": "number", "path": "moves.down.dead_end", "description": "yes/no" }, { "name": "down_tail_reachable", "type": "number", "path": "moves.down.tail_reachable", "description": "yes/no" }, { "name": "left_eats", "type": "number", "path": "moves.left.eats", "description": "yes/no" }, { "name": "left_food_distance", "type": "number", "path": "moves.left.food_distance" }, { "name": "left_room", "type": "number", "path": "moves.left.room" }, { "name": "left_dead_end", "type": "number", "path": "moves.left.dead_end", "description": "yes/no" }, { "name": "left_tail_reachable", "type": "number", "path": "moves.left.tail_reachable", "description": "yes/no" }, { "name": "right_eats", "type": "number", "path": "moves.right.eats", "description": "yes/no" }, { "name": "right_food_distance", "type": "number", "path": "moves.right.food_distance" }, { "name": "right_room", "type": "number", "path": "moves.right.room" }, { "name": "right_dead_end", "type": "number", "path": "moves.right.dead_end", "description": "yes/no" }, { "name": "right_tail_reachable", "type": "number", "path": "moves.right.tail_reachable", "description": "yes/no" } ] } } ``` ```bash curl https://api.canonopylabs.com/v1/domains -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d @snake.json ``` ### 3. Upload recorded moves Your cases are moments from recorded play, each with the move that was made: from a scripted player, a designer, or your best players. One JSON line per move: ```json {"state": {"heading": "right", "snakeLength": 3, "moves": {"up": {"eats": false, "food_distance": 12, "room": 253, "dead_end": false, "tail_reachable": true}, "...": "..."}}, "answers": {"move": "up"}} ``` ```bash curl https://api.canonopylabs.com/v1/domains/snake/data -H "Authorization: Bearer $CANONOPY_API_KEY" \ -F file=@recorded_moves.jsonl ``` ```json { "accepted": 3000, "rejected": 0, "total_cases": 3000, "per_question": { "move": 3000 }, "problems": [] } ``` ### 4. Train ```bash curl -X POST https://api.canonopylabs.com/v1/domains/snake/train -H "Authorization: Bearer $CANONOPY_API_KEY" ``` Ours took a few seconds and chose the recorded move on 99% of its held-back moves. The report also counts how often each rule applied; violations are always 0. See [Training and reports](https://canonopylabs.com/docs/training-and-reports.md#the-report). ### 5. Ask for each move Like Jev's demo, list only the legal moves as the question's options; the model picks among them. ```bash curl https://api.canonopylabs.com/v1/decide -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d '{ "model": "snake@latest", "state": { "heading": "right", "snakeLength": 3, "moves": { "up": { "eats": false, "food_distance": 12, "room": 253, "dead_end": false, "tail_reachable": true }, "right": { "eats": false, "food_distance": 12, "room": 253, "dead_end": false, "tail_reachable": true }, "down": { "eats": false, "food_distance": 14, "room": 253, "dead_end": false, "tail_reachable": true } } }, "questions": { "move": { "type": "choice", "instructions": "Which way should the snake move next?", "criteria": { "up": "move up", "right": "move right", "down": "move down" } } } }' ``` The move is `decision.action`. Leave out `questions` and it answers over all four directions; the rules still block the ones with no facts, and `decision.blocked_by_rules` lists them. In Python, with the SDK: ```python from canonopy import Client, Choice client = Client() # reads CANONOPY_API_KEY def next_move(state): legal = list(state["moves"]) res = client.decide( model="snake@latest", state=state, questions={"move": Choice("Which way should the snake move next?", {m: f"move {m}" for m in legal})}, ) return res.decision.action ``` ### 6. Run it inside the game loop A move every few milliseconds is where a download pays off: no network, nothing per move. ```bash curl -L -o snake.zip -H "Authorization: Bearer $CANONOPY_API_KEY" \ https://api.canonopylabs.com/v1/models/snake@latest/download ``` ```python from canonopy_runtime import load bot = load("snake.zip") def next_move(state): options = {m: f"move {m}" for m in state["moves"]} res = bot.decide({"state": state, "questions": {"move": {"type": "choice", "criteria": options}}}) return res["decision"]["action"] ``` The model is about 0.2 MB and reads numbers only, so it needs no text reader: `pip install https://canonopylabs.com/dl/canonopy_runtime-0.6.3-py3-none-any.whl`. See [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md). --- Source: https://canonopylabs.com/docs/cookbooks/ticket-triage (Markdown: https://canonopylabs.com/docs/cookbooks/ticket-triage.md) # Ticket triage, from day one A helpdesk with no history: route each ticket to billing, technical, account or sales, rate its priority, and flag customers who might leave. Here we start in day-one mode, with a general model answering from the first call, and let each question graduate. The other way to start with nothing is to [send the JSON you'd send Jev](https://canonopylabs.com/docs/start-with-no-data.md) and have your own model in minutes, as in the [zero-data bank cookbook](https://canonopylabs.com/docs/cookbooks/bank-zero-data.md). ## 1. Set it up Save this as `ticket-triage.json`: ```json { "name": "ticket-triage", "questions": { "team": { "type": "choice", "instructions": "Which team should handle this ticket?", "criteria": { "billing": "invoices, charges, plans and refunds", "technical": "bugs, errors, outages and integrations", "account": "login, access, users and settings", "sales": "pricing questions, upgrades and new contracts" } }, "priority": { "type": "score", "instructions": "How urgent is it?", "criteria": ["low", "normal", "high", "critical"] }, "churn_risk": { "type": "noul", "instructions": "Is this customer at risk of leaving?", "criteria": { "true": "they mention cancelling, switching or being fed up", "false": "they don't" } } }, "decision_question": "team", "rules": [{ "id": "no-sales-on-outage", "description": "A ticket about an outage never goes to sales", "when": [{ "field": "is_outage", "op": "is_true" }], "block": ["sales"] }], "state_schema": { "fields": [ { "name": "subject", "type": "text" }, { "name": "body", "type": "text" }, { "name": "plan", "type": "category", "values": ["free", "pro", "enterprise"] }, { "name": "seats", "type": "number" }, { "name": "is_outage", "type": "number", "description": "yes/no" } ] }, "graduation": { "min_outcomes": 3000, "auto": true } } ``` ```bash curl https://api.canonopylabs.com/v1/domains -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d @ticket-triage.json ``` ## 2. Pick a day-one backend A new domain starts with `local`, our built-in general model: no key, no cost. To use your own model instead, point it at any OpenAI-compatible endpoint: ```bash curl -X PUT https://api.canonopylabs.com/v1/domains/ticket-triage/fallback -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "provider": "openai", "base_url": "https://your-gateway.example/v1", "model": "your-model", "api_key": "" }' ``` Or Jev: `{ "provider": "typesafe", "api_key": "", "model": "jev-latest" }`. See [Day one](https://canonopylabs.com/docs/day-one.md#choose-your-backend). ## 3. Send tickets ```bash curl https://api.canonopylabs.com/v1/decide -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" -d '{ "model": "ticket-triage@latest", "state": { "subject": "API down", "body": "Our integration returns error 500 since this morning", "plan": "pro", "seats": 40, "is_outage": true } }' ``` Every answer comes back marked `"source": "fallback"`, with your rule applied to the `team` decision: ```json { "id": "dec_4c468a346f447e87", "model": "ticket-triage@day-one", "answers": { "team": { "type": "choice", "choice": "technical", "confidence": 0.93, "source": "fallback", "route": "act", "...": "..." }, "priority": { "type": "score", "score": 2.1, "confidence": 0.85, "source": "fallback", "route": "act", "...": "..." }, "churn_risk": { "type": "noul", "noul": 0.1, "source": "fallback", "route": "act" } }, "decision": { "question": "team", "action": "technical", "confidence": 0.93, "blocked_by_rules": ["no-sales-on-outage"], "blocked_actions": ["sales"], "route": "act" } } ``` A day-one answer below 0.8 confidence comes back `"route": "review"` and waits in the [unsure queue](https://canonopylabs.com/docs/improving.md#the-unsure-queue). ## 4. Report what happened When an agent closes a ticket, send the real team, using the `id` from the answer: ```bash curl https://api.canonopylabs.com/v1/outcomes -H "Authorization: Bearer $CANONOPY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "decision_id": "dec_4c468a346f447e87", "action": "technical", "answers": { "priority": 2, "churn_risk": false } }' ``` Or upload last quarter's export in one go: a CSV whose columns are `subject`, `body`, `plan`, `seats` and `team` (add `priority` and `churn_risk` columns if you have them). ```bash curl https://api.canonopylabs.com/v1/domains/ticket-triage/data -H "Authorization: Bearer $CANONOPY_API_KEY" \ -F file=@last_quarter.csv ``` ## 5. Watch it graduate ```bash curl https://api.canonopylabs.com/v1/domains/ticket-triage/graduation -H "Authorization: Bearer $CANONOPY_API_KEY" ``` ```json { "min_outcomes": 3000, "auto": true, "questions": { "team": { "status": "trained", "version": 1, "outcomes": 3202, "needed": 0 }, "priority": { "status": "fallback", "version": null, "outcomes": 1, "needed": 2999 }, "churn_risk": { "status": "fallback", "version": null, "outcomes": 1, "needed": 2999 } } } ``` At 3,000 outcomes a question trains and switches to your own model on its own, question by question: its answers change from `fallback` to `trained`. Graduation is checked each time an outcome or an unsure-queue answer comes in. After an upload that takes a question past 3,000, it graduates with the next outcome, or train it straight away: ```bash curl -X POST https://api.canonopylabs.com/v1/domains/ticket-triage/train -H "Authorization: Bearer $CANONOPY_API_KEY" ``` Once at least 20 tickets that your backend answered have a reported outcome, the training report compares the two on those tickets: `day_one` is your backend's accuracy there. If your backend is still more accurate on them, the question keeps using it for now. --- Source: https://canonopylabs.com/docs/api-reference (Markdown: https://canonopylabs.com/docs/api-reference.md) # API reference Base URL `https://api.canonopylabs.com`. Every request and answer is JSON. The machine-readable contract is OpenAPI 3.1; after this first version it only changes additively (new endpoints, new optional fields). **Auth:** `Authorization: Bearer ` (or `X-API-Key: `) on everything except `/health`, `/v1/plans`, `/v1/signup`, `/v1/login` and `/v1/demo/*`. Domain paths accept the domain's `name` or its `id`. ## Decide **`POST /v1/decide`** Alias: `POST /v1/systemone`, same body. - `model` (string): `@`, like `bank-support@latest`. Names starting with `jev-` map to your `default` domain. A domain that doesn't exist yet is created in day-one mode. - `state` (string | object | array, required): The case to decide. - `questions` (object): Choice, Score and Noul questions in Jev's format. Leave out to ask every question the domain knows. A Score with fewer than 2 levels is a `400` (Jev accepts it). Returns `id`, `model` (the version that answered), `answers` (each with `source`: `trained`, `translated`, `fallback`, `rule` or `backup`, and `route`: `act` or `review`), `language` (the language the text was detected in), `decision` (when the decision question was asked) and `usage` (`decisions`, plus `input_tokens` and `output_tokens` for Jev client compatibility; none of it is billed). See [Questions](https://canonopylabs.com/docs/questions.md) and [Decisions and rules](https://canonopylabs.com/docs/decisions-and-rules.md). **`GET /v1/models`** One card per serving domain version: `[{"name": "bank-support@3", "description": "…", "release_date": "2026-09-24"}]`. ## Build from a Jev request **`POST /v1/build`** Alias: `POST /v1/systemone/build`, same body. Your own model from the same JSON you send Jev. No past decisions or answers needed; a few examples of what you send Jev make it better. See [Start with no data](https://canonopylabs.com/docs/start-with-no-data.md). - `state` (string | object, required): ONE real example of the state your program sends. Each value is read as a field with a type and a `path`; identifiers, empty values and list items past the first 8 aren't read. - `questions` (object, required): Your questions exactly as you send them to Jev (Choice, Noul, Score). Stored as they are. - `examples` ((string | object)[]): Recommended: 5-20 more real states, each like `state` (up to 50; states only, no answers). - `fields` (object): What fields mean, keyed by a field's `path` or name: a sentence, or `{"meaning", "unit", "range": [lowest, highest]}`. E.g. `{"order.amount": "the refund asked for, in USD (0-5,000)"}`. - `rules` (string): Plain words. Rules that block answers of the decision question are enforced on every decision. - `name` (string): The model's name. Building again under the same name makes its next version. - `description` (string): What it decides. - `history` (object[]): Your past cases, up to 50,000: each `{"state", "answers"}` or a flat row, as in an upload. Stored as one upload of the model. See [Make it better with your history](https://canonopylabs.com/docs/with-your-history.md). - `interview` (boolean): `true`: [a few questions about your JSON](https://canonopylabs.com/docs/answer-a-few-questions.md) first; the build waits for your answers (status `awaiting_answers`) up to `answer_timeout_seconds` (default 600, at most 3600), then goes on with what it has. - `answer_timeout_seconds` (integer): With `interview`: how long it waits for answers. `model` and anything else Jev takes are ignored, so a pasted Jev request works as is. An old `sign_off` field is accepted and ignored. Returns a `Build` at once: - `build_id`, `domain`, `model` (`@latest`), `kind` (`build`, `retrain` or `rule_change`), `previous_build` (for a retrain or a rule change); - `status`: `awaiting_answers` (with `interview`: answer the questions), `queued`, `building`, `training`, `checking`, `ready` (serving), `needs_attention` (built, not serving: see `quality`) or `failed` (see `errors`). Nothing waits for you in between (a build made before Oct 1 can still show `awaiting_sign_off`; it moves on by itself). `phase` and `phases` (`understanding`, `questions` (with `interview`), `preparing`, `training`, `checking`, each with a `label` and a `status`); `message` and `next_step` in plain words; `notes`; - `understood`: `state` (`text` or `object`), `fields` (`name`, `path`, `type`), `not_read`, `questions`, `decision_question`, `rules`, `rules_not_understood`, `rules_note`, `field_notes` (`path`, `meaning`, `unit`, `range`) and `field_notes_not_matched` (`key`, `why`) when you sent `fields`, and `history` (`accepted`, `rejected`, `problems`) when you sent past cases; - `examples`: once `ready` or `needs_attention`, about 20 examples of how your model decides, optional to review (the same as `GET …/examples`); - `quality` once checked: `bar` (0.95), `passed`, `situations`, `statement`, and per question `agreement`, `passed`, `cases`, `advice` (below the bar) and `weakest` (`where`, `agreement`, `cases`); `own_cases` (`cases`, per-question `agreement`, `statement`) when your model was also measured on your own held-back cases; - `history`: `cases`, `learned_from`, `held_back`, `variations`, `flagged`, `pending`, `gaps` (`question`, `answer`, `your_cases`, `situations`), `statement`; - `found_rules`: rules found in your history, each `id`, `sentence`, `question`, `answer`, `support`, `agreement`, `statement`, `status` (`proposed`, `confirmed`, `rejected`), `hard`, `can_be_hard`, `changes` (`situations`, `of`, `statement`); - `flagged`: `total`, `pending` and the first 20 `items` (`id`, `state`, `answers`, `rule`, `question`, `rule_answer`, `kind`: `found` or `hard`, `status`: `pending`, `follow_rule`, `keep` or `drop`, `why`); - `questions`: the latest question round: `round`, `status` (`open`, `answered`, `timed_out`, `closed`), `count`, `answered`, `waits`, `asked_at`, `answer_by`, `answered_at`, `waited_seconds`, `used` (`this_build` or `next_retrain`); - `retrain` (a targeted retrain): `since_version`, `added` (`disagreements`, `around_confusions`, `look_alikes`, `example_messages`, `corrections`, `from_play`), `areas` (`question`, `answer`, `area`, `before`, `after`, `cases`), `statement`; - `rule_preview` (a rule change): `rules`, `previous_rules`, `situations`, `of`, `changes` (`question`, `from`, `to`, `situations`), about 10 `examples` (`state`, `before`, `after`), `statement`; - `stats` (`seconds`, `situations`, `examples`, `example_messages`, `download_bytes`, `model_file_bytes`, `version`), `version`, `errors`, `created_at`, `updated_at`, `finished_at`. `400` (the example or the questions can't be read), `402` (a build is a training for your plan), `409` (the name belongs to a model not made this way, or a build of it is in progress), `429` (a build limit: 5 builds per workspace per UTC day, 30 a month with a subscription, 3 in all on the free trial, where it's a `402` once billing is on; see [Build limits](https://canonopylabs.com/docs/errors-and-limits.md#build-limits)). **`GET /v1/build/{build_id}`** **`GET /v1/domains/{domain}/build`** The `Build`. The second is a model's latest build (its first build, a rebuild, a retrain or a rule change); `404` for a model not made with `POST /v1/build`. **`GET /v1/build/{build_id}/questions`** **`POST /v1/build/{build_id}/answers`** `GET` (`round`: 1 or 2, default the latest) → `build_id`, `domain`, `round`, `status`, `questions` (each `id`, `kind`, `text`, `priority`, `refs` (`question`, `choices`, `fields`), `answer_with`), `answered`, `rounds`, `asked_at`, `answer_by`, `answered_at`, `waited_seconds`, `next_step`. `POST {"answers": [{"id": "q02", "text": "…"}, {"id": "q05", "unit": "dollars", "range": [0, 5000]}, {"id": "q07", "skip": true}], "done": true}` → `build_id`, `round`, `accepted`, `rejected` (`id`, `why`), `used` (`this_build`, `next_retrain` or `none`), `status`, `message`. `400` (a key its question doesn't take, text over 600 characters). See [Answer a few questions](https://canonopylabs.com/docs/answer-a-few-questions.md). **`POST /v1/build/{build_id}/found-rules`** `{"rules": [{"id": "fr_1", "decision": "confirm", "hard": true}, {"id": "fr_2", "decision": "reject"}]}` → the `Build`. `decision`: `confirm` or `reject`; `hard` (with confirm, only when `can_be_hard`) also enforces it on every decision. Optional: nothing waits for it, and your decisions apply from the next retrain. `400` (an unknown id, or `hard` on a rule that can't be hard), `409` (the build is busy). See [Rules found in your history](https://canonopylabs.com/docs/rules-found.md). **`GET /v1/build/{build_id}/flagged?status=pending`** **`POST /v1/build/{build_id}/flagged`** `GET`: `status` (`pending` or `reviewed`; default all), `limit` (1-200, default 50), `cursor` → `build_id`, `domain`, `total`, `pending`, `items` (flagged past cases, as above), `next_cursor`. `POST {"rows": [{"id": "h1042", "decision": "follow_rule"}]}` or `{"all": "drop"}`: `follow_rule` (learn it with the rule's answer; not for a rule that only blocks answers), `keep` (learn it as it is) or `drop` (never learn it) → `build_id`, `domain`, `reviewed`, `total`, `pending`, `message`. Decisions apply when the model is next trained. See [Flagged past cases](https://canonopylabs.com/docs/rules-found.md#flagged-past-cases). ## Improve a model **`GET /v1/domains/{domain}/advice`** → `domain`, `statement` and `items`, most useful first: each `kind` (`retrain_to_sharpen`, `rule_may_be_wrong`, `answer_these`, `upload_these`, `from_playtest`, `shadow_review`, `answer_questions`), `title`, `text`, `action` (`retrain`, `change_rules`, `answer`, `upload`, `review`, `none`), and as they apply `question`, `agreement`, `rule`, `contradicted`, `applied`, `decisions` (`decision_id`, `summary`, `answer`, `confidence`). Read-only. See [Improve your model](https://canonopylabs.com/docs/improve-your-model.md#advice). **`POST /v1/domains/{domain}/rules`** For a model made with `POST /v1/build`: `rules` (plain words) and `mode` (`add`, the default, or `replace`) → a `Build` of kind `rule_change`, retrained with the new rules straight away. Its `rule_preview` says what changes, and its examples (the changed situations) are optional to review. The new version serves only if it passes the quality check. `400` (not a built model: use the set-up agent), `402`, `409` (a build is in progress), `429`. See [Change a rule with a preview](https://canonopylabs.com/docs/improve-your-model.md#change-a-rule-with-a-preview). **`POST /v1/domains/{domain}/decisions/{decision_id}/wrong`** **`POST /v1/decisions/{decision_id}/wrong`** `{"answers": {"action": "escalate"}}` or `{"action": "escalate", "note": "…"}` → `recorded`, `decision_id`, `domain`, `answers`, `message`. Recorded as the decision's outcome (`via: "marked_wrong"`); the next retrain learns it. `400` (an unknown question or answer), `404`. See [Mark a decision wrong](https://canonopylabs.com/docs/improve-your-model.md#mark-a-decision-wrong). A targeted retrain is `POST /v1/domains/{domain}/train` on a built model (below). ## Try it on real work **`PUT /v1/domains/{domain}/shadow`** **`GET /v1/domains/{domain}/shadow`** `PUT {"on": true}` (or `false`) → `domain`, `on`, `since`, `version`, `cases`, `agreed`, `disagreed`, `to_review`, `reviewed`, `agreement`, `statement`, `next_step`. See [Shadow mode](https://canonopylabs.com/docs/shadow-mode.md). **`POST /v1/domains/{domain}/shadow/cases`** `{"cases": [{"state": {…}, "current": {"action": "approve"}, "source": "jev", "confident": true}]}` (up to 100; `source`: `jev`, `process` or `person`) → `recorded`, `agreed`, `disagreed`, `items` (`index`, `id`, `agrees`, `error`), `message`. Your model decides each one silently: not acted on, not billed, not in the decision log. `409` (shadow mode is off, or no trained version yet). **`GET /v1/domains/{domain}/shadow/review?limit=10`** **`POST /v1/domains/{domain}/shadow/review`** `GET` (at most 20) → `domain`, `open`, `items` (`id`, `created_at`, `state`, `question`, `current`, `current_source`, `model`, `model_confidence`, `all_current`, `all_model`), `minutes`, `next_step`. `POST {"items": [{"id": "shd_…", "pick": "model"}, {"id": "shd_…", "answers": {"action": "escalate"}}, {"id": "shd_…", "skip": true}]}` → `reviewed`, `problems`, `open`, `message`. Reviewed answers are learned at your next retrain. **`GET /v1/domains/{domain}/spot-checks`** This week's optional card → `domain`, `setting` (`off`, `light`, `thorough`), `rate`, `title`, `open`, `minutes`, `items` (`decision_id`, `created_at`, `model`, `state`, `answers`, `decision`), `spot_checked` (`checked`, `agreed`, `accuracy`, `statement`), `next_step`. Answer each with `POST /v1/domains/{domain}/answers`. The setting is `PATCH /v1/settings {"spot_checks": "thorough"}`. **`POST /v1/domains/{domain}/play-results`** **`GET /v1/domains/{domain}/play-results`** `{"model": "my-bot@3", "episodes": [{"score": 1240, "deaths": 1}], "situations": [ … ]}` (1-1000 episodes of numbers; up to 300 optional situations) → `id`, `domain`, `version`, `episodes`, `summary` (per number: `mean`, `min`, `max`, `episodes`), `situations_saved`, `learned`, `statement`, `next_step`. `GET` (`limit` 1-50) → `domain`, `runs`. See [Playtest your model](https://canonopylabs.com/docs/playtest-your-model.md). ## Domains and set-up **`POST /v1/domains`** - `name` (string, required): Lowercase, used in `model`. - `description` (string): What it decides. - `kind` ("decision" | "game"): Defaults to `decision`. `game` is coming soon. - `questions` (object): Questions in Jev's format. - `decision_question` (string): The Choice question whose options are the actions. - `rules` (Rule[]): `id`, `description`, `when` (conditions on state fields), `block` or `allow_only`. - `state_schema` (object): `fields`: each `name`, `type` (`number`, `category`, `text`), optional `path`, `values`, `min`, `max`, `description`. - `graduation` (object): `min_outcomes` and `auto`. Returns the `Domain`: everything above plus `id`, `status` (`draft`, `ready`, `trained`), `actions`, `serving_version`, `latest_version`, `question_status` (per question: `trained` or `fallback`, `outcomes`, `needed`), `fallback`, `signoff`, `cases`, `created_at`, `updated_at`, and `safety_bar` (read-only: `target` 0.97, a plain-words `statement`, `strong_languages`, and per question the `threshold` and `automatic_share`). A request with `thresholds` gets a `400`: the bar sets itself. **`GET /v1/domains`** **`GET /v1/domains/{domain}`** **`PATCH /v1/domains/{domain}`** `PATCH` takes any subset of `description`, `questions`, `decision_question`, `rules`, `state_schema`, `graduation`, and returns the `Domain`. **`POST /v1/domains/{domain}/agent`** `{"message": "…"}` → `reply`, `summary` (the decision as set up so far, in plain words), `domain`, `open_questions`, `next_step` (`describe`: answer the open questions; `train`: it's set up, train it; `follow_build`: a rule change started; older set-ups can show `review_examples` or `upload_history`, both optional). On a model made with `POST /v1/build`, a message that changes the rules starts a rule change that retrains straight away, and the reply has its `build_id`. **`GET /v1/domains/{domain}/examples`** **`POST /v1/domains/{domain}/signoff`** Optional, any time: how your model decides on about 20 situations. `GET` → `examples` (each `id`, `state`, `answers`, `why`, `described`), `status` (`ready`, or `preparing` while the example messages of a decision made on text are being prepared: ask again in a minute), `note` and `signoff` (the review record; `status` `none` means not reviewed, which is fine). `POST {"examples": [{"id", "ok", "correct", "note"}], "retrain": true}` with only the examples you correct → the review record: `status` (`signed_off` or `corrections_recorded`), `at`, `by`, `examples`, `corrections`, and `message`. `retrain: true` starts the retrain in the same call: the answer has its `build_id` (a built model) or `job_id` (a described model). Without it, corrections apply at the next retrain. On a built model, corrections also go to your model's instructions, so the retrained model gives those answers. ## Past cases, training, versions **`POST /v1/domains/{domain}/data`** **`GET /v1/domains/{domain}/data`** Upload: a `.csv` or `.jsonl` file (multipart), a `text/csv` or `application/x-ndjson` body, or `{"rows": [...]}`. Query: `mode` (`add`, the default, or `replace`: these cases take the place of every uploaded case, in one step; nothing changes if no row is accepted) and `file_name` (names a raw body in the upload list). Returns `accepted`, `rejected`, `total_cases`, `per_question`, `problems` (first 20), `upload_id`, `mode` and `replaced` (`cases`, `uploads`). `GET` → `total_cases`, `per_question`, `from_uploads`, `from_signoff`, `from_outcomes`. **`GET /v1/domains/{domain}/uploads`** **`DELETE /v1/domains/{domain}/uploads/{upload_id}`** **`DELETE /v1/domains/{domain}/data?confirm={domain}`** `uploads` → `uploads` (newest first, each `id`, `source` (`upload`, or `signoff`: examples you corrected), `mode`, `file_name`, `created_at`, `cases`, `rejected`; cases from before uploads had ids are the group `earlier`), `total_cases`, `from_uploads`, `from_signoff`, `from_outcomes`. Deleting one upload, or every uploaded case (`confirm` must be the domain's name; add `signoff=true` and `outcomes=true` to also delete the examples you corrected and the hosted decisions the next version would learn from), deletes permanently and returns `removed` (`uploaded_cases`, `signoff_cases`, `decisions`), `uploads_removed`, `remaining` and a plain-words `message`. Versions already trained are kept. Both answer `409` while a new version of the domain is being made, as does `mode=replace`. See [Replace or remove cases](https://canonopylabs.com/docs/training-and-reports.md#replace-or-remove-cases). **`POST /v1/domains/{domain}/train`** Optional body: `questions` (only these), `promote` (`auto`, `always`, `never`), `wait` (default `true`). Returns a `Job`: `id`, `domain`, `kind` (`train`, `option`, `graduation`, `multilingual`, `starter` or `build`), `status` (`queued`, `running`, `done`, `failed`), `version`, `message`, `report`, and `build_id` for a model made with `POST /v1/build`. It learns from the cases the domain has now; the report's `cases` says how many, and how many earlier cases were removed since the last version. On a built model it's a [targeted retrain](https://canonopylabs.com/docs/improve-your-model.md#targeted-retrain): follow `GET /v1/build/{build_id}` for its `retrain` (what it added, and the before and after); with `wait: true` it waits for the build and answers with its training job. **`GET /v1/jobs/{job_id}`** **`POST /v1/domains/{domain}/options`** Adds a Choice option or an action: `question`, `option`, `description`, optional `when` (plain words), `never_when` (conditions; adds a rule), `also_allowed_under` (ids of `allow_only` rules that should also allow the new action), `promote`, `wait`. Returns a `Job` whose report has `before_after`. **`GET /v1/domains/{domain}/versions`** **`GET /v1/domains/{domain}/versions/{version}/report`** **`GET /v1/domains/{domain}/progress`** **`POST /v1/domains/{domain}/promote`** `versions` → `serving_version` and `versions` (each `version`, `status`, `kind`, `created_at`, `questions`, `recommendation`, `size_bytes`). `promote {"version": 2}` → `{"serving_version": 2}`; a roll-back is a promote of an older version. The report is described in [Training and reports](https://canonopylabs.com/docs/training-and-reports.md#the-report). `progress` → `waiting` (new cases since the latest version by `source`, `differed`, a plain `statement`, and `training` when a job is already running) and `versions`, newest first (each with `cases`, per-question `accuracy`, `change_points`, `comparison`: `same_cases`, `own_held_back` or `not_comparable`, `starter`, and `live` agreement with your reported outcomes from 20 of them). `waiting.sources` includes `marked_wrong` (decisions you marked wrong), and `advice` has the same items as `GET …/advice`. Read-only; see [Improving](https://canonopylabs.com/docs/improving.md#see-whats-waiting-then-retrain). **`GET /v1/models/{domain}@{version}/download`** `application/zip` with `model.json`, `heads.onnx` and `rules.json`. See [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md). **`GET /v1/readers/{reader}/{file}`** The shared text reader a downloaded model uses, by id (`reader-en-1`, about 34 MB, or `reader-multi-1`, about 113 MB: files `model.onnx` and `tokenizer.json`). `canonopy-runtime` fetches it once and checks it against the model's recorded checksum. ## Games (coming soon) **`PUT /v1/domains/{domain}/game`** (coming soon) **`GET /v1/domains/{domain}/game`** (coming soon) Game bots are coming soon; not yet available. Until they launch, these endpoints answer `503` with `"type": "unavailable"`. For `kind: "game"` domains. `PUT {"code", "episodes", "max_steps"}`: your game's Python, defining `initial_state(seed)`, `legal_moves(state)`, `step(state, move)` and optionally `features(state, move)`. It runs in a sandbox and is never returned. Returns `connected`, `moves`, `features`, `episodes`, `max_steps`, `checked_at`. Game bots are download-first: `/v1/decide` on a game domain answers `400`. To build a game player today, use a normal decision domain over the game state: see the [Snake](https://canonopylabs.com/docs/cookbooks/snake.md) and [Doom](https://canonopylabs.com/docs/cookbooks/doom.md) cookbooks. ## Outcomes, unsure queue, day one **`POST /v1/outcomes`** `decision_id`, `answers` (per question), `action` (shorthand for the decision question) → `recorded`, `decision_id`, `graduation`. **`GET /v1/domains/{domain}/unsure?limit=50`** **`POST /v1/domains/{domain}/answers`** `unsure` → `open` and `items` (each `decision_id`, `created_at`, `model`, `state`, `answers`, `decision`, `from_offline_device`). `answers {"decision_id", "answers"}` → `recorded`. **`POST /v1/domains/{domain}/offline-cases`** Cases a [downloaded model](https://canonopylabs.com/docs/running-it-yourself.md#save-and-sync-later) saved while it couldn't reach us; `model.sync()`, `canonopy-runtime sync` and `canonopy sync` send them for you. `{"cases": [...]}`, up to 100 per request (each as the runtime wrote it: `id`, `created_at`, `model`, `state`, `questions`, `answers`, `decision`) → `accepted`, `already_there` (sent before: nothing changed), `rejected` (each `id` and `reason`; they stay on the device), `decision_ids` and `held`. The held ones join the unsure queue with `from_offline_device: true`; outcomes take the runtime's `id` or the decision id. Answers from a device say `source` `trained`, `rule` or `local_fallback` (your own fallback). `402` after the trial with no subscription (the device keeps its cases), `409` for an archived model, `413` above 10 MB, `429` above 30 requests a minute. **`GET /v1/domains/{domain}/graduation`** **`GET /v1/domains/{domain}/fallback`** **`PUT /v1/domains/{domain}/fallback`** Fallback: `provider` (`typesafe`, `openai`, `local`, `none`), `api_key` (write-only), `base_url`, `model`. Reads return `has_key` and `key_last4`. ## Decision log **`GET /v1/domains/{domain}/decisions`** Newest first: `items` (each `id`, `created_at`, `model`, `version`, `summary`, `question`, `answer`, `confidence`, `route`, `source`, `sources`, `language`, `has_outcome`, `from_offline_device`, `marked_wrong`), `next_cursor` (pass it as `cursor` for older ones), `filters` and `retention_days`. Filters: `since`, `until` (ISO 8601; `until` is exclusive), `source`, `route`, `answer=question:answer` (repeatable), `version`, `has_outcome`, `q` (words in the case's text, or a decision id), and `limit` (1-200, default 50). **`GET /v1/domains/{domain}/decisions/{decision_id}`** One decision in full: the fields above plus `state`, `questions`, `answers`, `decision`, `rules_applied` (your rules that blocked an option, each `id`, `description`, `still_in_place`, and `chance` for a rule on what the message is about), `latency_ms` and `outcome` (`answers`, `reported_at`, `via`: `outcome`, `unsure_queue` or `marked_wrong`). **`GET /v1/domains/{domain}/decisions/export?format=csv`** The same filters, as CSV or JSONL (`format=jsonl`: one decision per line, in full), up to 100,000 decisions; `X-Decisions-Exported` and `X-Decisions-Truncated` say how many and whether there were more. **`DELETE /v1/domains/{domain}/decisions/{decision_id}`** **`DELETE /v1/domains/{domain}/decisions?before=2026-06-01&confirm={domain}`** Delete one decision, or every decision before a date (`confirm` is the model's name) → `deleted`, `learning_cases_deleted`, `message`. Permanent. `409` while a new version is being made if it would change what that version learns from. All of these keep working after the free trial ends. See [The decision log](https://canonopylabs.com/docs/decisions-and-rules.md#the-decision-log). ## Account **`GET /v1/account`** `org`, `user`, `plan`, and `usage`: `decisions_this_month`, `active_domains`, `monthly_price_usd` (always $20: the price is flat per workspace). **`GET /v1/keys`** **`POST /v1/keys`** **`DELETE /v1/keys/{key_id}`** `POST {"name": "production"}` returns the new key once, in `key`; we keep only a hash. **`GET /v1/settings`** **`PATCH /v1/settings`** `decision_retention`: how long decisions are kept, `30_days`, `90_days` (the default), `365_days` or `until_deleted`. Reads also return `decision_retention_days` and `older_than_retention` (deleted at the next daily clean-up). Older decisions are deleted once a day, including those with outcomes: later versions no longer learn from them. **`GET /v1/plans`** `[{"name", "price_usd_per_month": 20, "billing", "included", "rate_limit_per_second": 50}]`: $20 a month per workspace, flat, with unlimited decision models, decisions and retraining and up to 30 builds a month. (`price_usd_per_domain_month` is the same number, kept for older clients.) No key needed. **`POST /v1/signup`** **`POST /v1/login`** Sign-up (`org_name`, `email`, `password`; only when open sign-up is on) returns your first API key once. Login returns a 12-hour console session token, used as a Bearer key. **`POST /v1/billing/checkout`** Starts the subscription: one flat price, $20 a month for the workspace, however many decision models you have. Answers `503 unavailable` until billing is switched on. **`GET /v1/billing`** Your free trial and subscription: `trial` (`state`: `not_started` until your first trained model or first decision, then `active` for 5 days with `ends_at` and `days_left`, then `ended`), `trial_ends_at`, `status`, `can_train`, `can_decide`, `can_download` (always true), `grace_until` after a failed payment, and `message` when something is paused. After the trial with no active subscription, training, new options and `/v1/decide` answer `402 payment_required`; downloads never do. ## Open endpoints **`POST /v1/demo/decide`** **`GET /v1/demo/benchmark`** **`GET /health`** The demo decides with the `bank-support` model only, 30 requests per minute per IP. `/health` → `{"ok": true, "version": "0.1.0"}`. --- Source: https://canonopylabs.com/docs/cli-reference (Markdown: https://canonopylabs.com/docs/cli-reference.md) # CLI reference The `canonopy` command covers the whole workflow from a terminal: build a model from the JSON you send Jev (`canonopy build`), review how it decides (optional), add past cases, retrain, read reports and advice, decide, and manage versions and keys. It's optional: every command is one plain HTTP call to the [API](https://canonopylabs.com/docs/api-reference.md), and coding agents can use the [hosted MCP connector](https://canonopylabs.com/docs/coding-agents.md) with nothing installed. ## Install ```bash pip install https://canonopylabs.com/dl/canonopy_cli-0.6.2-py3-none-any.whl ``` The `canonopy-cli` package needs Python 3.9 or later and nothing else. `pipx install` works with the same URL. Set your key from the [console](https://canonopylabs.com/console/keys): ```bash export CANONOPY_API_KEY=cnp_… ``` - `--key` (option): Your API key. Default: `$CANONOPY_API_KEY`. - `--base-url` (option): The API. Default: `$CANONOPY_BASE_URL`, or `https://api.canonopylabs.com`. - `--json` (option): Print the API's JSON instead of a summary. ## Commands In the order you'd usually use them. A decision model is called a *domain* in the API. ```bash canonopy build jev-request.json --name refunds --rules "Never approve when amount is over 500." # your own model from a Jev request canonopy build status refunds # where the build is (or: a build id) canonopy build jev-request.json --interview # a few questions about your JSON first canonopy build questions BUILD_ID # the questions (or: a model's name) canonopy build answer BUILD_ID # answer them on the terminal (or --file answers.json) canonopy playtest my-bot@latest --cmd "python my_game.py" --episodes 20 # your game plays with the model, locally canonopy shadow start refunds # shadow mode: beside your current decisions canonopy shadow send refunds requests.jsonl # real requests with your current decision canonopy shadow review refunds # the disagreements (--pick-model, --pick-current, --skip) canonopy spot-checks refunds # this week's optional card (--set off|light|thorough) canonopy init refunds --describe "We handle refund requests. Actions: approve, escalate, decline. …" canonopy describe refunds "Card vs payment: these go to cards" # the set-up conversation (- reads stdin) canonopy examples refunds # optional: how it decides on about 20 cases canonopy signoff refunds --correct ex_03:action=escalate --retrain # optional: correct any, retrain with them canonopy upload refunds past_cases.csv # CSV with a header row, JSONL or JSON canonopy upload refunds better_cases.csv --replace # these replace every uploaded case canonopy uploads refunds # each upload: id, date, file, cases canonopy uploads delete refunds UPLOAD_ID # one upload's cases, deleted permanently canonopy data clear --confirm refunds # every uploaded case (--signoff, --outcomes) canonopy train refunds # --no-wait, --promote auto|always|never canonopy advice refunds # what would improve it most canonopy report refunds # or: canonopy report refunds 3 canonopy decide refunds@latest --state '{"message": "refund please", "amount": 40}' canonopy options refunds action store-credit --description "offer store credit" --when "tier is pro" canonopy outcomes DECISION_ID --action approve # what really happened canonopy outcomes --unsure refunds # the unsure queue canonopy outcomes DECISION_ID --review refunds --answer action=escalate canonopy progress refunds # what's waiting for the next retrain, and every version canonopy versions refunds canonopy promote refunds 3 # what refunds@latest serves canonopy download refunds@latest -o refunds.zip # to run yourself canonopy fallback refunds --provider typesafe --key-env TYPESAFE_API_KEY canonopy keys [list | create NAME | revoke KEY_ID] canonopy decisions refunds --route review --since 2026-09-01 # the decision log (--json for JSON) canonopy decisions show refunds DECISION_ID # one decision in full canonopy decisions export refunds --format csv > decisions.csv # or --format jsonl, -o FILE canonopy decisions delete refunds DECISION_ID # or --before 2026-06-01 --confirm refunds canonopy decisions wrong DECISION_ID --answer escalate # mark a decision wrong (--note "…") canonopy retention [30_days | 90_days | 365_days | until_deleted] # how long decisions are kept canonopy sync # cases a downloaded model saved offline canonopy sync --status # where they are, and how many wait canonopy mcp # the local MCP server canonopy skill install # the agent skill for Claude Code (--project, --dir, --force) ``` `--state` and `--questions` take JSON, plain text, or `@file.json`. Errors print the API's message and exit with status 1. `canonopy sync` sends what a [downloaded model](https://canonopylabs.com/docs/running-it-yourself.md#save-and-sync-later) saved while it couldn't reach us (it needs the `canonopy-runtime` package installed alongside; `--dir` picks another folder). Each case leaves the local file once it's stored; the held ones join your unsure queue. `canonopy decisions` filters with `--since`, `--until`, `--source`, `--route`, `--answer action:refund` (repeatable), `--version`, `--has-outcome` or `--no-outcome`, and `--search` (words in the case's text, or a decision id); `--limit` and `--cursor` page through. Export takes the same filters. See [The decision log](https://canonopylabs.com/docs/decisions-and-rules.md#the-decision-log). `upload --replace`, `uploads delete` and `data clear` delete cases permanently; examples you corrected and outcomes stay unless `data clear` gets `--signoff` or `--outcomes`, and versions already trained are kept. Train again afterwards for the next version of the same model. See [Replace or remove cases](https://canonopylabs.com/docs/training-and-reports.md#replace-or-remove-cases). ## Build from a Jev request For models made from the JSON you send Jev ([Start with no data](https://canonopylabs.com/docs/start-with-no-data.md)), in `canonopy-cli` 0.4.0 and later: ```bash canonopy build jev-request.json # the body you send Jev, as is canonopy build jev-request.json --name refunds --description "Refund requests" \ --rules "Never approve when amount is over 500. Always escalate when fraud_flag is true." canonopy build jev-request.json --name refunds --examples more-states.jsonl \ --field "order.amount=the refund asked for, in USD (0-5,000)" # 5-20 more real states, notes on fields canonopy build jev-request.json --name refunds --history past_refunds.csv --wait canonopy build status refunds # the latest build of a model, or a build id canonopy examples refunds # optional, once it's ready: how it decides canonopy signoff refunds --correct ex_03:action=escalate --retrain # optional: correct any, retrain with them canonopy rules found refunds # rules found in your history canonopy rules found refunds --confirm fr_3 --hard fr_1 --reject fr_2 canonopy history flagged refunds --status pending # past cases that contradict your rules canonopy history flagged refunds --follow-rule h1042 --keep h1180 --drop h2210 canonopy history flagged refunds --all follow_rule # or keep, or drop canonopy advice refunds # what would improve it most canonopy train refunds # a targeted retrain: prints the build id canonopy rules change refunds "Always escalate when amount is over 300." # retrains now; --replace for all your rules canonopy decisions wrong DECISION_ID --answer escalate # or --answer action=escalate, --note "…" ``` - `build FILE` (command): `FILE` holds `{"state", "questions"}`, and optionally `examples`, `fields`, `rules` and `name` (a pasted Jev request works: `model` is ignored). `--examples FILE` (5-20 more real states, like `state`: one per line, or a JSON list; up to 50; no past decisions or answers needed, a few examples of what you send Jev make it better), `--fields FILE` (notes on fields: a JSON object keyed by path or name) and `--field "path=note"` (repeatable), `--rules "…"` or `--rules-file F` (plain words), `--name`, `--description`, `--history FILE` (past cases: CSV, JSONL or JSON), `--wait` (until the model is ready, needs attention or failed; nothing waits for you in between), `--json`. - `build … --interview` (command): The build asks [a few questions about your JSON](https://canonopylabs.com/docs/answer-a-few-questions.md) first and waits for your answers up to `--answer-timeout` seconds (default 600), then goes on with what it has. With `--wait`, it returns when the questions are ready. - `build questions` (command): A build id or a model's name (its latest build): the questions, most useful first, and what each answer holds. `--round 2` for the follow-up. - `build answer` (command): Asks each question on the terminal (Enter skips it), or sends `--file answers.json` (`{"answers": [{"id": "q01", "text": "…"}]}`). `--more`: more answers follow, the build keeps waiting. - `build status` (command): A build id (`bld_…`) or a model's name: where the build is, what it read, the quality check, and what to do next. - `rules found` (command): Lists the rules found in your history with their support, agreement and what confirming each changes. `--confirm`, `--reject` and `--hard` take rule ids (`fr_1 …`); `--hard` confirms a rule and enforces it on every decision, like your own rules. A build id works in place of the model's name. - `history flagged` (command): Lists flagged past cases (`--status pending` or `reviewed`, `--limit`, `--cursor`). `--follow-rule`, `--keep` and `--drop` take case ids; `--all` decides for every pending case. - `rules change` (command): New rules in plain words (or `-` to read them from stdin), added to yours (or `--replace`). Your model is retrained with them straight away, and the build says what changed; the new version serves only if it passes the quality check. `--wait` waits for the build to finish. Its examples (the changed situations) are optional to review with `canonopy examples`. - `decisions wrong` (command): The right answer for a decision: `--answer escalate` (the decision question) or `--answer question=answer` (once per question), and `--note`. `canonopy decisions wrong MODEL DECISION_ID` works too. The next retrain learns it. - `advice` (command): What would improve the model most, in plain words. - `examples` (command): Optional, any time: how your model decides on about 20 situations, with the answer for each and why. - `signoff` (command): Optional: `--correct EXAMPLE_ID:question=value` for each example you correct (repeatable), and `--retrain` to retrain with them now. Without `--retrain`, they apply at the next retrain. `--all-ok` also marks every other example right. On a model made with `canonopy build`, `canonopy train` starts a targeted retrain and prints its build id; `canonopy build status` shows the before and after. See [Rules found in your history](https://canonopylabs.com/docs/rules-found.md) and [Improve your model](https://canonopylabs.com/docs/improve-your-model.md). ## Try it on real work ```bash canonopy playtest my-bot@latest --cmd "python my_game.py" --episodes 20 [--once] [--keep-situations 200] canonopy shadow start|stop|status refunds canonopy shadow send refunds requests.jsonl canonopy shadow review refunds [--pick-model ID … --pick-current ID … --skip ID …] canonopy spot-checks refunds | canonopy spot-checks --set thorough ``` - `playtest` (command): Serves the model (a local `.zip`, or `name@version` downloaded from your workspace) at a local endpoint on 127.0.0.1, runs `--cmd` once per episode (or once for all with `--once`), reads one JSON line of numbers per episode, and sends only those numbers (`--no-send` to keep them). Needs `canonopy-runtime`. See [Playtest your model](https://canonopylabs.com/docs/playtest-your-model.md). - `shadow` (command): `start`, `stop`, `status`; `send MODEL FILE` (JSON lines of `{"state", "current", "source", "confident"}`); `review` lists the disagreements, and `--pick-model`, `--pick-current`, `--skip` record your review. See [Shadow mode](https://canonopylabs.com/docs/shadow-mode.md). - `spot-checks` (command): This week's optional card for a model, or `--set off|light|thorough` for the workspace. ## The local MCP server `canonopy mcp` runs an MCP server on stdio so coding agents can use every step above, plus a `docs` tool. It reads your key from `CANONOPY_API_KEY` and never prints it. ```bash claude mcp add canonopy --env CANONOPY_API_KEY=cnp_… -- canonopy mcp ``` Most agents don't need it: the [hosted MCP connector](https://canonopylabs.com/docs/coding-agents.md) has the same tools with nothing installed. Use the local server to upload files straight from disk, save downloaded models to disk, or read a day-one backend's key from an environment variable. Setup for Cursor, Codex and other clients: [Use with coding agents](https://canonopylabs.com/docs/coding-agents.md#the-local-mcp-server). ## The agent skill `canonopy skill install` copies the [Canonopy Decisions skill](https://canonopylabs.com/docs/coding-agents.md#skill) to Claude Code's personal skills folder, `~/.claude/skills/canonopy-decisions/`, and prints where it went. `--project` installs it in the current project's `.claude/skills/` instead (or `--project DIR`), and `--dir PATH` in any agent's skills folder (as `PATH/canonopy-decisions/`). It never replaces an installed copy unless you add `--force`. `canonopy skill show` prints `SKILL.md`. ## The runtime CLI The runtime that runs a downloaded model offline has a command line of its own, in the `canonopy-runtime` package: ```bash pip install "canonopy-runtime[text] @ https://canonopylabs.com/dl/canonopy_runtime-0.6.3-py3-none-any.whl" python -m canonopy_runtime bank-support@3.zip request.json canonopy-runtime sync # send the cases it saved while offline (reads CANONOPY_API_KEY) canonopy-runtime queue # where they are, and how many wait canonopy-runtime playtest my-bot@3.zip --cmd "python my_game.py" --episodes 20 # your game plays with it; prints the results canonopy-runtime reader list # the text readers downloaded to this machine, and their size canonopy-runtime reader remove # delete them to free the space (fetched again when a model needs one) ``` `request.json` is a `/v1/decide` body. The answer is printed in the same format as the hosted endpoint. `sync` and `queue` take `--dir FOLDER` (default: `CANONOPY_OFFLINE_DIR` or your user data folder); `sync` exits with status 1 when it couldn't send everything (the cases stay in the file). See [Running it yourself](https://canonopylabs.com/docs/running-it-yourself.md#save-and-sync-later). ## Or with curl ```bash export CANONOPY_API_KEY=cnp_… API=https://api.canonopylabs.com curl -s $API/v1/domains -H "Authorization: Bearer $CANONOPY_API_KEY" # list domains curl -s $API/v1/domains/refunds/data -H "Authorization: Bearer $CANONOPY_API_KEY" -F file=@cases.csv # upload curl -s -X POST $API/v1/domains/refunds/train -H "Authorization: Bearer $CANONOPY_API_KEY" # train curl -s $API/v1/domains/refunds/versions -H "Authorization: Bearer $CANONOPY_API_KEY" # versions ``` --- Source: https://canonopylabs.com/docs/errors-and-limits (Markdown: https://canonopylabs.com/docs/errors-and-limits.md) # Errors and limits ## Errors Errors always look like this, with a short, plain message: ```json { "error": { "type": "invalid_request", "message": "Send `state`: text, a JSON object or an array." } } ``` | Status | `error.type` | What to do | |---|---|---| | 400 | `invalid_request` | fix the request; the message says what's wrong | | 401 | `authentication_error` | check the key and the `Authorization` header | | 402 | `payment_required` | the free trial has ended and there's no active subscription, or the trial's 3 builds are used: subscribe in the console (Billing). The models you already have are yours to download | | 404 | `not_found` | check the domain name, version or id | | 409 | `conflict` | it already exists, or the domain isn't ready for that yet | | 413 | `too_large` | split the upload | | 429 | `rate_limited` | wait for `Retry-After` (1 second for the rate limit), then retry; a build limit says when it opens again | | 500 | `server_error` | retry with backoff; tell us if it persists | | 503 | `unavailable` | a feature that isn't switched on yet, or a short outage | ## Rate limits The hosted endpoint allows **50 requests per second per key** by default. Above that you get `429` with `Retry-After: 1`. Nothing is ever charged for it: pricing is flat, $20 a month per workspace, with unlimited decision models, decisions and retraining. ## Build limits A build is a new model, a new version from `POST /v1/build`, or a rule change. Each workspace can start: | | Builds | |---|---| | a day (UTC) | 5 | | during the free trial, in all | 3 | | a calendar month, with a subscription | 30 | Retraining a built model doesn't count as a build; it has a fair-use limit of 20 retrains a day. Past a limit you get `429 rate_limited` (or `402 payment_required` when the trial's builds are used up), with a message that says which limit was reached and when it opens again. Your models keep working, and you can still retrain them. `GET /v1/account` shows where you are in `usage.builds`, and the console shows it under Billing. ## Free trial and subscription Your 5-day free trial starts when your workspace trains its first model or makes its first decision, not when you sign up, and needs no card. After it, the subscription ($20 a month per workspace, flat) pays for improving your models: training, new options and actions, day-one questions switching to your own model, and the hosted endpoint. Without one, those answer `402 payment_required` (and automatic switching waits); everything else keeps working: the console, reports and versions, uploads, set-up, reviewing examples and billing. Your models are yours to keep. Every version you trained, during the trial or while subscribed, can be downloaded at any time, and a downloaded model keeps working offline. `GET /v1/billing` says where you are: `trial.state` (`not_started`, `active` or `ended`), `trial_ends_at`, and `can_train`, `can_decide` and `can_download`. Need more throughput? Ask several questions in one request (they're answered together), or [run the model yourself](https://canonopylabs.com/docs/running-it-yourself.md) for any volume. The open demo endpoint (`/v1/demo/decide`) allows 30 requests per minute per IP. ## Retries - Retry `429`, `500` and `503` with exponential backoff and jitter. - Don't retry `400`, `401`, `402`, `404`, `409` or `413` unchanged. - A retried `/v1/decide` is answered again and counted as another decision; with flat pricing that costs nothing. ## Sizes - Choice questions take 1 to 255 options; Score questions 2 to 10 levels. - Uploads through the console go up to 20 MB; send bigger files straight to the API. - `train` with `wait: true` waits up to about 10 minutes; use `wait: false` and poll `GET /v1/jobs/{job_id}` for longer runs.