Unofficial guide · Jev / TypeSafe
Jev confidence thresholds: how to choose act, review and escalate bands
There is no universal Jev confidence threshold. TypeSafe's own guidance is to split answers into bands (act, proceed with caution, don't act), set stricter cutoffs for riskier actions, and tune the numbers on your own data.
Independent guide. aiedu.guide is not affiliated with TypeSafe AI or Jev, and is not endorsed or sponsored by either. Threshold numbers quoted here come from public vendor examples and are illustrations, not recommendations.
Every threshold number on this page is quoted from an example in vendor docs. None is a recommended default.
Our Jev vs LLM guide gives the short answer (label your own cases and tune the cutoff on them). This page is the full method.
What is Jev confidence, and how is it different from probability?
Short answer: For a Choice or a Score, confidence is a number from 0 to 1 that summarizes how concentrated the probability distribution is. It is not the chance that the answer is right, and in TypeSafe's API reference a Noul has no confidence field: its noul value is the probability of yes.
TypeSafe describes confidence as "a statistic computed from the probability distribution" that "collapses that shape into a single number from 0 to 1". A flat distribution gives low confidence. A peaked one gives high confidence. TypeSafe has not published the exact formula, and says it will add a separate cookbook on the computation later.
TypeSafe's launch blog says "All answers are accompanied with calibrated probabilities and confidence scores". Its API reference and Confidence page are narrower: confidence comes on Choice and Score answers, and "Noul answers don't carry one".
| Signal | Question types | What it tells you |
|---|---|---|
noul | Noul | The probability that the answer is yes. It is both the answer and the certainty. |
An option's value in probabilities | Choice, Score | How much weight that one option or level gets. TypeSafe's API reference says the values sum to 1; the AI SDK docs note that after rounding they may not add up to exactly 1 (see Rounding below). |
confidence | Choice, Score | How concentrated the whole distribution is |
score | Score | A probability-weighted position on your levels (0 to n-1), not a percentage |
Why confidence is not the winner's probability. TypeSafe's classification cookbook gives the example: a winner at 0.45 with a runner-up at 0.44 is a different situation from a winner at 0.45 with the rest of the weight scattered thinly. The top probability is the same. The confidence is not.
What calibration does and doesn't promise. TypeSafe trains Jev with RLCD (Reinforcement Learning for Calibrated Decisions) and says its probabilities are "optimized against outcomes". Calibration means outcomes given 0.2 "should occur about 20% of the time", and outcomes given 0.8 about 80%. TypeSafe adds: "These rates describe groups of predictions, not a guarantee about any single answer." Its launch post claims "higher confidence means higher accuracy". We found no published calibration metrics in TypeSafe's docs.
Confidence is not correctness. TypeSafe's Score page says a confidence of 1.0 "describes the model's answer, not a guarantee that the answer is correct".
Is Jev's confidence the probability that the answer is correct? No. It summarizes how concentrated the Choice or Score distribution is. Only labeled outcomes from your own workflow tell you how it relates to errors.
What confidence threshold should I use for Jev?
Short answer: No single number. TypeSafe says "A confidence threshold is not one number": the right cutoff depends on your domain, on how costly each kind of mistake is, and on how the model performs on your own labeled examples.
TypeSafe's full caveat: "The correct threshold values depend on your domain and the performance of the model for your use case. Start with conservative thresholds, test with your own data, and adjust". OpenRouter and Vercel say the same:
- OpenRouter: "pick thresholds from the cost of each kind of mistake rather than from a round number".
- Vercel: "Calibrate probabilities and confidence against labeled examples from your workflow." The AI SDK docs add: "Choose decision thresholds in application code."
Vendor examples do use numbers. Here they are in one place, so you can see how much they vary.
Thresholds used in published examples (illustrations, not defaults)
| Where | What is gated | Numbers used in the example | Context |
|---|---|---|---|
| TypeSafe: Confidence page | confidence | Below 0.5 routes to a human; above 0.9 before a high-stakes action | Thresholds scale with risk |
| TypeSafe: Confidence-gated routing pattern | confidence | Floor of 0.6; a transfer auto-approved only above 0.85; 0.6 to 0.85 asks the user to confirm | Approving a transfer |
| TypeSafe: Noul page | noul | 0.5 when yes and no are equally easy to act on; code examples use > 0.9, and a YES 0.8 / NO 0.2 band with the middle sent to review | Code examples |
| TypeSafe: Noul self-consistency cookbook | noul | Uncertain band 0.30 to 0.70, "neither a calibrated guarantee nor an optimized threshold" | Repeatability test |
| TypeSafe: Choice self-consistency cookbook | Top value in probabilities (not confidence) | 0.60 or above, otherwise "uncertain" | Repeatability test |
| TypeSafe: Classification using confidence cookbook | confidence | CONFIDENT = 0.9: report the specific group, otherwise the broader division | Falling back to a broader label |
| TypeSafe: LLM guardrails cookbook | Guardrail policy decisions | Review at 0.35; act at 0.70 (strict) or 0.85 (permissive); severity block at 2.0 | Guardrail policies |
Two TypeSafe pages alone use different numbers (0.5 and 0.9 on the Confidence page, 0.6 and 0.85 in the routing pattern), which shows the values depend on context. Both self-consistency cookbooks say to set production thresholds from labeled examples and from the cost of errors and of review. The guardrails cookbook says to "set the thresholds ... from labeled examples of your own traffic".
How do I set act, review and escalate bands?
Short answer: Work out what each wrong action costs, pin the model version, run Jev on labeled examples from your own traffic, and compare accuracy per confidence band. Then set cutoffs where the errors in the "act" band are ones you can accept, and adjust as you learn.
TypeSafe's suggested pattern has three bands:
| Band | TypeSafe's label | What happens |
|---|---|---|
| High | "Act automatically" | Code acts on the answer |
| Medium | "Proceed with caution" | Proceed carefully, for example by asking the user to confirm (as in TypeSafe's routing pattern) |
| Low | "Do not act" | Route to a human, ask a clarifying question, or fall back |
The steps below are our editorial method for filling in those bands. They follow the vendor guidance quoted above, but the wording is ours.
- List the actions and their costs. For each action the answer can trigger, write down what a wrong one costs. A mis-sorted queue item is cheap to undo; a wrong refund is not.
- Pin the model version. Use
jev-1.13.0on TypeSafe ortypesafe/jev-1.13on OpenRouter. See pin jev-1.13.0 vs jev-latest and where confidence lives on each Jev channel. - Build a labeled set from your own traffic. Label each case independently of Jev, and include the ambiguous ones.
- Run Jev on it and log everything. Keep the full
probabilities,confidence(ornoul), and the responsemodel. - Group results by band and compare with your labels. Record how many cases fall in each band as well as the error rate, so a small band doesn't look better than it is.
- Choose cutoffs. Trade the errors you accept in the "act" band against the review work the other bands create.
- Start conservative, then adjust. This is TypeSafe's own advice. Loosen a cutoff only when your data supports it.
- Re-run the set when anything changes: the model version, the question wording, or the kind of traffic you send.
If you track this in a spreadsheet, these columns cover it: input_id, human_label, jev_answer, top_probability, confidence, response_model, band, auto_action_correct (Y/N), sent_to_review.
How should thresholds differ for Noul, Choice and Score?
Short answer: Threshold a Noul's noul value directly, ideally with separate yes and no cutoffs and a review band between them. Gate a Choice on confidence per action, optionally with the selected option's probability, and treat a Score as a position on your levels, not a percentage.
For background on the three types, see pick the right question type (Noul, Choice or Score).
Noul thresholds
TypeSafe's Noul page: "Use 0.5 when yes and no are equally easy to act on." Raise the cutoff when a false yes is expensive, lower it when a missed yes is expensive, and send middle values to a person.
Does a low Noul (boolean) probability mean Jev is unsure? No. A value near 0 means Jev puts a low probability on yes. For a Noul, yes and no are most evenly balanced at 0.5.
Does a Noul have a confidence value? No, not in TypeSafe's API reference. The noul value is itself the probability of yes. Threshold it directly, ideally with separate yes and no cutoffs.
Choice thresholds
confidence and the selected option's value in probabilities measure different things (see the 0.45 / 0.44 example above), so decide which one each rule gates on and keep it consistent. Add an other option so unclear inputs have somewhere to go. When confidence is low, you don't have to stop: TypeSafe's classification cookbook reports the broader division instead of the specific group when confidence is below its example cutoff.
Score thresholds
Put the cutoff on score, which runs on your level numbers (0 to n-1), and check confidence alongside it. Don't read score as a percentage. TypeSafe's jaggedness page says jev-1.13 score levels "are weak in numerical calibration", so don't interpolate exact magnitudes from a score.
Should thresholds depend on how risky the action is?
Short answer: Yes. TypeSafe says thresholds should scale with risk: use one low floor for "don't act at all", and a higher bar for high-stakes actions such as approving a transfer.
TypeSafe's confidence-gated routing pattern shows the idea. Its example numbers, as an illustration only:
- Confidence below 0.6: don't act on the answer.
- A transfer with confidence above 0.85: approve automatically.
- A transfer with confidence from 0.6 to 0.85: ask the user to confirm.
Should I use the same threshold for every action? No. Use one low floor for "don't act at all", and a separate, higher bar for each action that is hard to undo. It's the shape of TypeSafe's example that carries over, not its numbers.
Why do my Jev thresholds break after a model or channel change?
Short answer: Aliases move to new versions, Noul and Choice numbers aren't interchangeable, and each channel exposes confidence differently. Re-run your labeled set before reusing any cutoff.
Check these:
- The alias moved.
jev-latestpoints to a new release when one ships. TypeSafe: "If you have tuned confidence thresholds against a specific version, pin that version's ID instead of the alias". Log the responsemodel. - A threshold moved between question types. TypeSafe's jaggedness page shows the same question returning 0.22 as a Noul and 0.01 for "yes" as a Choice. Don't carry a Noul threshold over to a Choice.
- The channel moved confidence. In the AI SDK, Choice and Score confidence sits in
result.providerMetadata.typesafe.confidence, not on the answer. If it's missing, send the case to your fallback rather than guessing a value. - Rounding. The AI SDK docs say TypeSafe rounds scores and probabilities to 2 decimals, so probabilities may not sum to exactly 1. Don't set cutoffs finer than two decimals.
- The wording changed. New criteria wording can shift the distribution. Re-run the labeled set.
- The language changed. English is Jev's primary training language. For other languages, TypeSafe says to test first and "pay close attention to Confidence when routing".
Do I need to re-tune when Jev updates? If you use an alias like jev-latest, yes: aliases move to new versions. Pin a versioned ID such as jev-1.13.0 and re-run your labeled set before you switch.
Field paths for every channel are in where confidence lives on each Jev channel.
What should happen when Jev fails, times out or returns no confidence?
Short answer: Treat it like TypeSafe's low band: don't act, and send the case to a defined fallback such as a human queue or an LLM. Count those cases when you measure error rates.
TypeSafe's SDKs retry 429 and 529 responses with backoff, and the Python SDK's default timeout is 10.0 seconds. After that, your code decides. A missing value is not a low value: route it, don't score it. For a route-first design that hands off to an LLM, see route first, verify after: using Jev with an LLM.
When is changing the threshold the wrong fix?
Short answer: When accuracy on your labeled set doesn't improve as confidence rises. Then the question needs fixing (split it, reword it, sharpen the criteria, or trim the state), or the task isn't a good fit for Jev.
Moving a cutoff can't fix a question that mixes two conditions or a Choice without an other option. Start with pick the right question type (Noul, Choice or Score), and check when Jev is the wrong tool.
FAQ
What is a good default confidence threshold for Jev?
No single number. TypeSafe says "A confidence threshold is not one number": the right cutoff depends on your domain, on how costly each kind of mistake is, and on how the model performs on your own labeled examples.
Is Jev's confidence the probability that the answer is correct?
No. It summarizes how concentrated the Choice or Score distribution is. Only labeled outcomes from your own workflow tell you how it relates to errors.
Does a Noul have a confidence value?
No, not in TypeSafe's API reference. The noul value is itself the probability of yes. Threshold it directly, ideally with separate yes and no cutoffs.
Does a low Noul (boolean) probability mean Jev is unsure?
No. A value near 0 means Jev puts a low probability on yes. For a Noul, yes and no are most evenly balanced at 0.5.
Should I use the same threshold for every action?
No. Use one low floor for "don't act at all", and a separate, higher bar for each action that is hard to undo. It's the shape of TypeSafe's example that carries over, not its numbers.
Do I need to re-tune when Jev updates?
If you use an alias like jev-latest, yes: aliases move to new versions. Pin a versioned ID such as jev-1.13.0 and re-run your labeled set before you switch.
Next steps
Official sources
Facts on this page come from these pages. Jev launched in September 2026 and details change quickly, so check the official page before you rely on a number.
- TypeSafe docs: Confidence
- TypeSafe docs: Confidence-gated routing
- TypeSafe docs: Noul
- TypeSafe docs: Score
- TypeSafe docs: Models (aliases, pinning, languages)
- TypeSafe docs: System One (calibration, RLCD)
- TypeSafe docs: Machine learning primer (what calibration means)
- TypeSafe docs: Jev 1.13 jaggedness (last reviewed 17 Sep 2026)
- TypeSafe docs: API reference (errors and retries)
- TypeSafe docs: Python SDK constants (default timeout)
- TypeSafe cookbook: Classification using confidence
- TypeSafe cookbook: Self-consistency with Noul
- TypeSafe cookbook: Self-consistency with Choice
- TypeSafe cookbook: LLM guardrails
- TypeSafe blog: Introducing System One Models & Jev (15 Sep 2026)
- OpenRouter docs: Jev tutorial
- Vercel changelog: TypeSafe AI's Jev now available on AI Gateway (16 Sep 2026)
- AI SDK docs: TypeSafe AI provider
Checked 25 September 2026