Skip to content
Docs / Confidence

Confidence

Every Choice and Score answer has a confidence; a Noul is a probability itself. They're calibrated: across many answers where it says 90%, it's right about 90% of the time.

Measured on your cases, not promised

Calibration is measured across groups of answers, not guaranteed for any single one. Every training report shows yours, on your own held-back cases:

  • a calibration error (0 is perfect; the bank test scored 0.006, against 0.044–0.063 for Jev);
  • a reliability curve: stated confidence against how often it was right;
  • a sentence you can quote, like "When it says 90–100% sure, it's right 99.3% of the time."

The safety bar sets itself

You don't choose a cut-off. After every training run, the model sets its own safety bar: the lowest confidence at which its automatic answers are at least 97% right on your held-back cases, checked language by language. Each retrain measures it again, and the report says it in plain words: "automatic answers 97%+ accurate; 78% handled automatically; the rest double-checked".

Where the safety bar sits

A real curve: 3,080 held-back bank messages from our benchmark. Each bar is the share of messages at or above a confidence. The model sets its own bar at the lowest confidence where automatic answers are at least 97% right.

The bar
0.60
Handled automatically
95.9%
Of those, right
97.5%

From 0.6: 95.9% of messages, 97.5% right. Automatic.

Below the bar, or in a language the model isn't strong in, a message is double-checked: translated to English and decided again, then your day-one backend, else held for a person. See Languages.

Using confidence in code

python
d = res["decision"]
if d["route"] == "act":
    run(d["action"])               # automatic: at or above the safety bar
else:
    hand_to_person(res["id"])      # review or ask: it's also in your unsure queue
  • Prefer route over raw numbers: it follows the model's safety bar, which moves with every retrain. Code written for Jev that reads confidence keeps working.
  • Check source. A fallback answer's confidence comes from your day-one backend, calibrated its own way; a backup answer's comes from our hosted backup.
  • Middling scores need a look at confidence. A Score of 1.5 can mean "clearly in the middle" or "torn between the ends". See the explorer on Questions.