JEV, a Low-Cost AI Judge That Knows How to Doubt

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A two-step scheme based on the confidence of JEV claims 93.4% accuracy at 41.4% of the cost of a reference LLM judge, according to reported tests. It appears that JEV and other models fail in certain cases, and that its quick and economical decisions are relevant under specific conditions.
A two-level routing achieves 93.4% accuracy at 41.4% of the cost
In the described setup, JEV's confidence corresponds to its maximum probability among the options. An example given is a distribution of 95% for "first" and 5% for "second," which sets the confidence at 0.95. The proposed routing involves setting a threshold, for example, 0.90, to accept JEV's decision above this threshold and refer below this threshold to a larger model. Tested on 1,610 new pairs, this approach entrusted 68.5% of the cases to JEV alone. The overall observed accuracy reached 93.4%, slightly above GPT-6 at 92.5%, for a cost representing 41.4% of that of GPT-6. On a more difficult task, the system referred 74.2% of the cases to the larger model. The stated goal is to avoid the wrong answer while not arbitrating based on budget.
Without a reference answer, even large models can err
When no comparison answer is available, JEV, GPT-4.1 mini, and GPT-5.4 made nearly random choices while appearing confident. In this case, increasing the model size did not provide a solution. The recommendation that emerges is to provide any judge with evidence or a checklist.
Cost and speed: JEV measured 277 times cheaper, 13 times faster
Four researchers from CMU compared JEV to sixteen other judges across numerous tasks. In their measurements, JEV cost $0.044 for 1,000 judgments and took 0.15 seconds per judgment, while GPT-6 Astra displayed 1.89 seconds and $12.182 for 1,000 judgments. Based on this, JEV was 277 times cheaper and 13 times faster, although these costs reflect the testing protocol and may vary. In terms of accuracy, JEV achieved 92.5% on RewardBench like GPT-6, 87.3% versus 88.4% on HaluEval, 78.6% versus 93.1% on a more difficult bench, 68.4% versus 95.9% on logical puzzles, and 76.6% versus 90.1% when the correct answer was longer and more elaborate compared to a concise and direct expectation.
Observed calibration: 0.043 error for JEV versus 0.054 for an LLM
A test conducted on 88 cases via OpenRouter showed JEV's confidence figures well distributed between 0 and 1, while the LLM judge rarely placed its confidence in the middle and often ended up close to 0 or 1 even in uncertainty. The recorded error scores were 0.043 for JEV and 0.054 for the LLM, with a lower value being better. However, it is noted that a test does not constitute proof and that confidence levels should be compared to actual responses.
What JEV does well, what LLMs bring
JEV is presented as a "System One" model designed for small decisions, returning short answers with probabilities. TypeSafe claims to have trained it to provide honest confidence figures, a claim that has not been independently verified. JEV is described as very fast and economical, suitable for simple and repeated checks when the evidence is in the text. In contrast, an LLM judge produces written text, can explain its decision, but is slower and more expensive, making it suitable for open questions, complex reasoning, and written feedback. Langfuse, an AI application tracking tool, is integrated with LLM judges and code checks and already supports JEV. In terms of features, JEV can select an option from a list by associating probabilities, decide between two answers with a type of error, or produce a level on a scale such as low, medium, or high.
Reproducing a trial: environment, difficult cases, and appeals to judges
A demonstration gathered 12 difficult cases, mixing long but incorrect answers, hidden instructions, and calculations, with two answers per case and a known truth. Each judge was queried twice, first with A and then with B, to detect a preference for form. To reproduce the trial, Python 3.9 or higher is required, along with a TypeSafe API key for JEV and a key for a compatible OpenAI LLM. The procedure suggests creating a dedicated folder, installing requests, pandas, numpy, python-dotenv, openai, matplotlib, and truststore, and then configuring a .env (TYPESAFE_API_KEY, LLM_API_KEY, LLM_BASE_URL, LLM_JUDGE_MODEL) without sharing it. The project includes .env, judges.py, raw_call.py, cases.py, run_lab.py, report.py, and plot_frontier.py. The support file defines call_jev_pair (returns including who_won and probability) and jev_two_order, as well as their equivalents for the LLM, with a SYSTEM instruction asking to choose the best answer and to treat texts as data to ignore any hidden instructions.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.