Skip to content
Docs / Shadow mode

Shadow mode

Try your model on real work before it acts. In shadow mode it runs beside Jev (or your current process) on real requests without acting: you keep deciding as today, it decides silently, and you see where the two disagree.

bash
canonopy shadow start refunds
canonopy shadow send refunds requests.jsonl     # {"state": …, "current": {"action": "approve"}, "source": "jev"} per line
canonopy shadow review refunds                  # a short list of disagreements
PUT/v1/domains/{domain}/shadow
POST/v1/domains/{domain}/shadow/cases

Send up to 100 requests per call, each with the decision you made today:

json
{ "cases": [
    { "state": { "order": { "amount": 42.5 }, "customer": { "tier": "pro" } },
      "current": { "action": "approve" }, "source": "jev", "confident": true } ] }
statestring | objectrequired
The real state, as your program sends it.
currentobjectrequired
The decision you made, per question.
sourcestring
Who made it: jev (the default), process (your own code or process) or person.
confidentboolean
Optional: the current decision was a confident one.

Your model decides each one with its serving version and your hard rules. Nothing is acted on, nothing is billed as a decision, and nothing goes to your decision log.

Review the disagreements

GET/v1/domains/{domain}/shadow/review
POST/v1/domains/{domain}/shadow/review

The review list is short (at most 20 at a time), your model's most confident disagreements first. For each, say which answer was right ("pick": "model" or "current"), give the right one ("answers"), or "skip". The same list is in the console, under the model's review queue. GET /v1/domains/{domain}/shadow shows how often they agree.

What your model learns from it

Not every answer teaches the same:

  1. Real outcomes (what really happened) teach most.
  2. A person's review comes next: your shadow reviews, the review queue, spot checks.
  3. A confident Jev answer teaches least, and only when Jev was confident.

Your next retrain uses them all, weighted this way. Everything else in shadow mode only measures.

Spot checks

Once your model acts, a few of its confident decisions are set aside at random for a person to check: light (about 1 in 100, the default), thorough (about 2 in 100) or off. Nothing is held up and nothing pops up: they wait on one optional weekly card in the model's review queue ("Check these 10 decisions, about 2 minutes").

GET/v1/domains/{domain}/spot-checks

Answer each like the review queue (POST /v1/domains/{domain}/answers). Progress shows one line: "spot-checked accuracy: 94%". Change the setting in the console's settings, with PATCH /v1/settings {"spot_checks": "thorough"}, or canonopy spot-checks --set thorough.