Skip to content
Docs / Bank support, day one

Bank support: start in minutes, then improve with your decisions

A bank's support inbox: each message is about one of eleven topics and goes to one of nine next steps, under five hard rules. This cookbook builds the desk's model from the request it would send Jev, with no past messages, then shows how it pulls ahead as it learns from the desk's decisions, on 3,080 real customer messages.

The result

3,080 real messagesAccuracy
Day one, built from the Jev request with no data (10 automatic builds)85.5% (84.6–86.8%)
After about 400 reviewed cases (the cases it was least sure about, answered)91.7%
After about 800 reviewed cases92.7%
With the bank's history (about 9,000 past messages)about 95% (94.9–96.0%)
Jev, few-shot (22 example messages in the prompt)86.6%
Jev, zero-shot (the written policy only)83.9%

Close to Jev on day one; clearly ahead once it learns from your decisions. A build takes about 5 minutes. On day one it's level with few-shot Jev, within a point or two either way, and ahead of zero-shot Jev. Answer the cases it sends to review and retrain: after about 400 of them it's 5 points ahead of Jev. If you'd rather just report outcomes as they come, about 800 random cases get it to about 90%, and 1,600 to about 92.6%.

The model is about 830 KB to download and shares the standard 34 MB English text reader.

Read these before quoting the numbers

  • The messages are the banking77 test split (PolyAI, CC BY 4.0). It's public, so either model may have seen it before.
  • Accuracy is against a written routing policy, the same for both sides. The account details attached to each message were assigned at random for the test; only the text is real.
  • Day one: the 9 builds were made by POST /v1/build on Sep 30 with slightly different settings while we tuned it. Two more builds with a bug since fixed, and two experiments with other settings, scored 81.6–86.75% and aren't in the range.
  • The reviewed cases came from the bank's real past messages (never the test ones), answered with their true answers, as a person reviewing them would; each step is the mean of 2 runs. That test started from a no-data model made from hand-written example messages (88.6% before any cases), a stronger start than today's automatic builds. Answers that come mostly from a backup model, not a person, will likely help less.
  • Below about 200 cases, don't expect a visible change: steps of 25 to 100 cases move it less than the ±1 point between two trainings.
  • Rules on your fields (amounts, verification, the fraud flag) are never broken. Rules on what a message is about, like "block a lost card first", are enforced whenever there's a real chance they apply, and unsure cases go to review; on day one some messages about lost cards are read as something else, so a few of those cases can be missed. Fewer are missed as it learns; with the bank's full history and those rules, the test had none.
  • Long messages are different: on real consumer complaints (CFPB narratives, a few hundred words each), a model built with no data scored about 40% against Jev's 63%. With the complaints' history it scored 77.7% against Jev's 65.0% few-shot. For long texts, start from your past cases.

Build it

1. Take the request you'd send Jev

One real example of the state (the message and the account fields), and the two questions with a one-line description of each option. Save it as bank-jev-request.json:

json
{
  "model": "jev-latest",
  "state": { "message": "I think someone has my card, there are payments I did not make",
             "tier": "plus", "account_age_days": 812, "amount": 86.5, "prior_disputes": 0, "fraud_flag": false,
             "card_status": "active", "identity_verified": true, "contacts_7d": 1, "new_device_login_24h": false },
  "questions": {
    "topic": { "type": "choice", "instructions": "What is the customer writing about?",
      "criteria": {
        "lost_stolen": "lost, stolen or compromised card (or a lost or stolen phone with the app on it)",
        "unrecognised": "a payment, cash withdrawal or direct debit they don't recognise, a double charge, or the ATM gave the wrong amount of cash",
        "refund": "they want a refund from a merchant, or a refund hasn't shown up",
        "card_fault": "their card doesn't work: declined, contactless or virtual card failing, PIN blocked, or the ATM swallowed it",
        "card_delivery": "getting, ordering, activating, linking or replacing a card, card arrival, changing the PIN",
        "payments": "a bank transfer or payment is pending, failed, declined, reverted, cancelled or hasn't arrived, or how to send or receive money",
        "topup": "topping up the account: a top-up is pending, failed or reverted, top-up methods and limits, cash or cheque deposits not showing",
        "fees": "a fee or charge they were billed (card payment, transfer, top-up, cash withdrawal, or an extra charge on the statement)",
        "fx": "exchange rates, currency exchange, a wrong exchange rate, or which currencies are supported",
        "identity": "verifying their identity or source of funds, a forgotten passcode, or editing personal details",
        "info": "general product questions: where the card is accepted, Visa or Mastercard, supported countries, ATMs, Apple Pay or Google Pay, card limits, age limits, closing the account" } },
    "action": { "type": "choice", "instructions": "What should the support desk do?",
      "criteria": {
        "block-card": "block the customer's card straight away so nobody can use it",
        "refund": "refund the money to the customer",
        "open-dispute": "open a chargeback dispute on the transaction",
        "escalate-fraud": "hand the case to the fraud team",
        "route-payments": "send the case to the payments and transfers team",
        "route-cards": "send the case to the cards team",
        "request-verification": "ask the customer to verify their identity first",
        "answer-faq": "answer the question directly from the help centre",
        "escalate-senior": "hand the case to a senior human agent" } }
  }
}

The option descriptions matter most here: they're how your model knows what each topic covers.

2. Build, with the desk's policy in plain words

Save the desk's policy, one rule per sentence, as policy.txt:

text
Never refund when amount is over 250.
Never refund when identity_verified is false.
When fraud_flag is true, only escalate-fraud or block-card.
Always block-card when topic is lost_stolen and card_status is active or frozen.
Never answer-faq when topic is unrecognised.
escalate-fraud when fraud_flag is true.
route-cards when topic is lost_stolen.
request-verification when topic is identity.
request-verification when identity_verified is false and topic is one of refund, unrecognised, payments, topup.
escalate-senior when topic is unrecognised and prior_disputes is at least 3.
open-dispute when topic is unrecognised.
answer-faq when topic is fees and tier is not one of premium, metal.
escalate-senior when topic is refund or fees and amount is over 250.
escalate-senior when topic is refund or fees and account_age_days is under 30.
refund when topic is refund or fees.
escalate-senior when topic is one of card_fault, card_delivery, payments, topup and contacts_7d is at least 3.
route-cards when topic is card_fault or card_delivery.
route-payments when topic is payments or topup.
Otherwise answer-faq.
bash
curl https://api.canonopylabs.com/v1/build -H "Authorization: Bearer $CANONOPY_API_KEY" \
  -H "Content-Type: application/json" \
  -d "$(jq --rawfile rules policy.txt '. + {name: "bank-support", rules: $rules}' bank-jev-request.json)"

Or canonopy build bank-jev-request.json --name bank-support --rules "$(cat policy.txt)" --wait.

In understood:

  • fields: message (text) and the nine account fields: tier and card_status as categories, fraud_flag, identity_verified and new_device_login_24h as yes/no, the rest as numbers.
  • decision_question: action.
  • rules: the first five sentences, which block answers, are enforced on every decision. The first three read account fields, so they hold whatever the model thinks. The two about topic depend on what the message is about: they're enforced whenever there's a real chance they apply, and unsure cases go to review. See Rules on what a message is about.
  • The routing sentences ("route-cards when topic is card_fault or card_delivery") aren't hard rules; your model follows them in its answers.

3. Wait for ready, then switch over

There's nothing to do in between: canonopy build status bank-support shows where it is. Once it's ready, change the base URL and model, and keep sending the same JSON:

diff
- POST https://api.typesafe.ai/v1/systemone      "model": "jev-latest"
+ POST https://api.canonopylabs.com/v1/systemone  "model": "bank-support@latest"

Each answer says where it came from and whether to act on it; decision.blocked_by_rules says which rules stopped which actions. A message it's unsure about comes back with route: "review": send those to a person, and their answers teach the next retrain the most.

4. See how it decides (optional)

Any time once it's ready: about 20 messages with their account details, each with the topic and action your model gives, and why.

bash
canonopy examples bank-support
canonopy signoff bank-support --correct ex_05:topic=unrecognised --retrain   # only if one is wrong

Make it better with your decisions

This is where it pulls ahead. Two ways, and you can do both.

Answer the cases it was unsure about. Every message it isn't sure of comes back with route: "review" and waits in the unsure queue. Answer them in the console (Improve) or the API, then retrain; GET /v1/domains/bank-support/advice says which cases help most. Those are the reviewed cases above: about 400 of them took it to 91.7%. See Improve your model.

Send your history. The desk's past messages, each with the topic or action it got. Build again under the same name:

bash
canonopy build bank-jev-request.json --name bank-support --rules "$(cat policy.txt)" --history past_messages.csv --wait
  • Your past messages become the main examples, and your model is measured on some it never saw: the report says how it does on your own cases.
  • Rules found in your history are proposed for you to confirm, each with how many past cases it covers. Only the ones you confirm are used. See Rules found in your history.
  • Past messages that contradict your rules are set aside for you to review, never learned silently.

With the bank's history, the same kind of model scored about 95% (94.9–96.0%). The numbers, and building from history alone: Bank support with history.