j.jevmanual.
OfficialChecked 2026-09-21·jev-1.13.0

Why Is My Jev Confidence Low?

Ten causes of low confidence, with bad and better examples you can test before changing thresholds.

On this pageDiagnose before lowering the gate1. Overlapping Choice options2. Missing Other option3. Vague instructions4. Insufficient state5. Too much irrelevant state6. Ambiguous rubric7. Task requires more reasoning8. Contradictory criteria9. Wrong primitive10. Language differencesWhat to measure after a fix

Diagnose before lowering the gate

Low confidence can be a useful signal that the question, evidence, or answer space is ambiguous. Start by preserving the input, model version, question revision, and probability distribution. A valid response with uncertainty is different from an API error.

1. Overlapping Choice options

Bad: billing and payments as separate unexplained categories.

Better: Define distinct responsibilities or merge the overlap.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

2. Missing Other option

Bad: Force every message into three departments.

Better: Add an explicitly described other category for unmatched work.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

3. Vague instructions

Bad: “Classify this”.

Better: Ask which team owns ticket.message and why each option fits.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

4. Insufficient state

Bad: Rate refund eligibility without the relevant policy.

Better: Include verified payment facts and the applicable policy excerpt.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

5. Too much irrelevant state

Bad: Send the entire customer history.

Better: Filter to the relevant message and current account facts.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

6. Ambiguous rubric

Bad: “Low / medium / high” with no definitions.

Better: Define an ordered operational rubric with recognizable boundaries.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

7. Task requires more reasoning

Bad: Ask for a multi-hop numerical proof.

Better: Use code or a suitable reasoning model, then classify a simpler judgment.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

8. Contradictory criteria

Bad: Instructions say urgent; levels describe customer value.

Better: Align the question with the rubric’s single dimension.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

9. Wrong primitive

Bad: Ask which category using Noul.

Better: Use Choice for categories, Score for levels, Noul for one proposition.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

10. Language differences

Bad: Assume an English-tested threshold works unchanged in every language.

Better: Evaluate each language and route unsupported confidence bands to review.

Test the revised question against borderline examples as well as obvious ones; improving certainty alone does not prove better accuracy.

What to measure after a fix

Compare labeled accuracy, false automatic actions, and review volume. Keep a held-out evaluation set so changes do not merely fit your debugging examples. A model can be confidently wrong; sample automatic outcomes as well as low-confidence cases.

Search the manual