Skip to content
Docs / Bank support, with history

Bank support with history, head to head with Jev

A bank's support inbox: each message goes to one of nine next steps, under five hard rules. We ran the same 3,080 real customer messages through hosted Jev and through a model learned from the bank's cases. This cookbook shows the numbers, then how to build a model like it from your history.

No history yet? Build the same decision from its Jev request in minutes: close to Jev on day one (85.5% against 86.6%), and 91.7% after about 400 reviewed cases. See Bank support, from day one. Your history takes it further.

The result

CanonopyJev, few-shotJev, zero-shot
Accuracy96.0% ± 0.486.6%83.9%
Calibration error0.0060.0440.063
Time per decision (p50)38.9 ms on a laptop CPU; 0.49 ms once the text is read177 ms over HTTPS216 ms over HTTPS
Cost per million decisions$0.78 compute$158.61 at list price$55.58 at list price

How much history this took. Learned from history alone, with no zero-data start: from 1,000 cases it scored 78.3%, below Jev; from 3,000, 88.2%; from 9,000, 94.4%; from 27,000, 96.0%. A case is one message with its account details; the 27,000 are about 9,000 real messages, each with three account states. Starting from your Jev request instead, you don't wait for that history: the model built with no data is close to Jev from day one, and every case you review makes it better from there (91.7% after about 400). See Bank support, from day one.

Read these before quoting the numbers

  • The messages are the banking77 test split (PolyAI, CC BY 4.0). It's public, so either model may have seen it before.
  • Accuracy is against a written routing policy. Some of Jev's misses were defensible routing calls; counting those as right, the lead is still about 5–7 points.
  • The account details attached to each message were assigned at random for the test. Only the text is real; real account data is still untested.
  • Our times and costs were measured on a laptop CPU; Jev's are HTTPS wall time and list price.

Build it

This builds a model like it from your own history: past messages with their account details and what the desk did.

1. Describe the decision

bash
curl https://api.canonopylabs.com/v1/domains -H "Authorization: Bearer $CANONOPY_API_KEY" \
  -H "Content-Type: application/json" -d '{"name": "bank-support"}'

curl https://api.canonopylabs.com/v1/domains/bank-support/agent -H "Authorization: Bearer $CANONOPY_API_KEY" \
  -H "Content-Type: application/json" -d '{"message": "We route messages for a bank support desk. Actions: block-card, refund, open-dispute, escalate-fraud, route-payments, route-cards, request-verification, answer-faq, escalate-senior. Fields: message (text), tier (category: basic, plus, premium, metal), amount (number), fraud_flag (yes/no), card_status (category: active, frozen, blocked), identity_verified (yes/no), prior_disputes (number), contacts_7d (number), account_age_days (number). Never refund when amount is over 250; escalate-senior instead. Never refund when identity_verified is false; request-verification instead. When fraud_flag is true, only escalate-fraud or block-card. Otherwise answer-faq."}'

The reply: Set up 9 actions, 9 state fields and 3 rules that can't be broken. Next, train it (looking at the example cases is optional, any time). Those three rules read account fields, so they're enforced on every decision. The benchmark's other two rules (block a lost card before anything else; never answer an unrecognised payment from the FAQ) depend on what the message is about. Give the domain a topic question and say them as rules ("Always block-card when topic is lost_stolen and card_status is active or frozen. Never answer-faq when topic is unrecognised.") and they're enforced whenever there's a real chance they apply, with unsure cases sent to review. See Decisions and rules.

2. Upload your history

One row per past message: its account fields, and the action the desk took, in a column named action.

message,tier,amount,fraud_flag,card_status,identity_verified,prior_disputes,contacts_7d,account_age_days,action
"I lost my card on the bus",plus,40,0,active,1,0,1,812,block-card
"I was charged twice for the same coffee",premium,312.4,0,active,1,0,0,1460,escalate-senior
bash
curl https://api.canonopylabs.com/v1/domains/bank-support/data -H "Authorization: Bearer $CANONOPY_API_KEY" -F file=@tickets.csv

3. Train

bash
curl -X POST https://api.canonopylabs.com/v1/domains/bank-support/train -H "Authorization: Bearer $CANONOPY_API_KEY"

Optional, any time: look at how it decides on 20 example cases. Since you uploaded first, they're your own messages. Open Review examples in the console, or use GET /v1/domains/bank-support/examples; to correct any, POST /v1/domains/bank-support/signoff with "retrain": true (see Training and reports).

4. Call both, side by side

python
import os, requests

ACTIONS = {
    "block-card": "block the card straight away", "refund": "refund the money",
    "open-dispute": "open a dispute with the merchant", "escalate-fraud": "hand to the fraud team",
    "route-payments": "send to the payments team", "route-cards": "send to the cards team",
    "request-verification": "ask the customer to verify their identity",
    "answer-faq": "answer from the help centre", "escalate-senior": "hand to a senior agent",
}
state = {"message": "I was charged twice for the same coffee", "tier": "premium", "amount": 312.4,
         "fraud_flag": False, "identity_verified": True}
questions = {"action": {"type": "choice", "instructions": "What should the support desk do?", "criteria": ACTIONS}}

ours = requests.post("https://api.canonopylabs.com/v1/systemone",
    headers={"Authorization": f"Bearer {os.environ['CANONOPY_API_KEY']}"},
    json={"model": "bank-support@latest", "state": state, "questions": questions}).json()
jev = requests.post("https://api.typesafe.ai/v1/systemone",
    headers={"Authorization": f"Bearer {os.environ['TYPESAFE_API_KEY']}"},
    json={"model": "jev-latest", "state": state, "questions": questions}).json()

print(ours["answers"]["action"]["choice"], ours["decision"]["blocked_by_rules"])
print(jev["answers"]["action"]["choice"])

Same request, same answer shape. Ours adds source and route on each answer and the decision block. Here the amount is over 250, so the first rule fires: blocked_by_rules is ["never-refund-when-amount-is-over-250-escalate-senior-instead"], and refund can't be the answer whatever the model thinks. The rule's id comes from its wording; GET /v1/domains/bank-support lists them.

5. Read the report

Each training report shows accuracy on your held-back messages, the weakest actions, the pairs it confuses most with real messages, and a ranked list of what would help. In our benchmark the weakest topic was lost or stolen cards (89%), and most mistakes were between answer-faq, route-cards and route-payments, on questions a reasonable reader could route either way.