Skip to content
Docs / Languages

Languages

Every model reads about 100 languages; accuracy is measured in 51. A customer can write in Spanish, Hindi or Japanese, and the same model decides, with your rules enforced.

Strong languages

Each version knows which languages it's strong in. Its training report lists them, with results per language, and so does the domain (safety_bar.strong_languages). A strong language is one where the model's automatic answers clear the safety bar on your held-back cases.

The safety bar: 97%, set for you

Automatic answers are held to a fixed bar: at least 97% right. After every training run, the model sets the confidence each answer needs to reach it, from your own held-back cases, and it's checked language by language. Each retrain measures it again. There's no threshold to tune.

Unsure messages are double-checked

Two checks run in code on every answer, never left to the model:

  • Language: a message in a language the model isn't strong in is always double-checked.
  • Confidence: an answer below the bar is double-checked.

A double-checked message is translated to English and the same model decides again, with your rules still applied. That answer says "source": "translated". If it's still below the bar, your day-one backend (Jev or an LLM) answers ("source": "fallback"). Without one, our hosted backup answers from your allowed options, with your rules still applied ("source": "backup"). If the backup is unavailable or can't tell, it's held for a person: "route": "review", and it waits in your unsure queue.

json
{
  "answers": {
    "action": { "type": "choice", "choice": "block-card", "confidence": 0.98,
                "probabilities": { "block-card": 0.98, "open-dispute": 0.02 },
                "source": "translated", "route": "act" }
  },
  "decision": { "question": "action", "action": "block-card", "confidence": 0.98,
                "blocked_by_rules": [], "blocked_actions": [], "route": "act" },
  "language": "es"
}

language is the language the message was written in. It's null when there's too little text to tell (a number, or "ok"): the confidence check still applies.

A weak language gets better with cases

When a language is below the bar, the training report says so and suggests what fixes it, for example "Upload ~300 cases in Swahili to handle more automatically". Past cases written in that language, with how they were resolved, raise the share it handles on its own.

Running it yourself

A downloaded model runs the same two checks. With your API key set, it hands a double-checked message to us. Offline, it holds the message for review ("route": "review"). It never guesses.