Confidence
Every Choice and Score answer has a confidence; a Noul is a probability itself. They're calibrated: across many answers where it says 90%, it's right about 90% of the time.
Measured on your cases, not promised
Calibration is measured across groups of answers, not guaranteed for any single one. Every training report shows yours, on your own held-back cases:
- a calibration error (0 is perfect; the bank test scored 0.006, against 0.044–0.063 for Jev);
- a reliability curve: stated confidence against how often it was right;
- a sentence you can quote, like "When it says 90–100% sure, it's right 99.3% of the time."
The safety bar sets itself
You don't choose a cut-off. After every training run, the model sets its own safety bar: the lowest confidence at which its automatic answers are at least 97% right on your held-back cases, checked language by language. Each retrain measures it again, and the report says it in plain words: "automatic answers 97%+ accurate; 78% handled automatically; the rest double-checked".
Where the safety bar sits
A real curve: 3,080 held-back bank messages from our benchmark. Each bar is the share of messages at or above a confidence. The model sets its own bar at the lowest confidence where automatic answers are at least 97% right.
- The bar
- 0.60
- Handled automatically
- 95.9%
- Of those, right
- 97.5%
From 0.6: 95.9% of messages, 97.5% right. Automatic.
Below the bar, or in a language the model isn't strong in, a message is double-checked: translated to English and decided again, then your day-one backend, else held for a person. See Languages.
Using confidence in code
d = res["decision"]
if d["route"] == "act":
run(d["action"]) # automatic: at or above the safety bar
else:
hand_to_person(res["id"]) # review or ask: it's also in your unsure queue- Prefer
routeover raw numbers: it follows the model's safety bar, which moves with every retrain. Code written for Jev that readsconfidencekeeps working. - Check
source. Afallbackanswer's confidence comes from your day-one backend, calibrated its own way; abackupanswer's comes from our hosted backup. - Middling scores need a look at
confidence. A Score of 1.5 can mean "clearly in the middle" or "torn between the ends". See the explorer on Questions.