A few weeks ago a new kind of model arrived with a loud claim. TypeSafe released Jev, which it calls a "decision model": it cannot write a paragraph, but it can answer yes-or-no questions about your data. The launch coverage led with it being hundreds of times cheaper and faster than a large language model.
We did not take that at face value. But our first reaction was simple: if even half of it is true, it changes how everyone should think about AI evaluation, guardrails and automated decisions.
So we evaluated it ourselves. Here is what we tried, how we evaluated it, what we found, and what we have now built into AgentGuard.
The sampling problem
Every team running AI in production wants to know one thing: are the answers right? The standard way to find out is an "LLM judge": a second AI model that reads each answer and grades it.
The catch is cost. A judge has to read everything your app read (the question, the documents, the answer) and then write out its reasoning. So each check is roughly as big a job as the call that produced the answer. When the AI app is already a large line on the bill, adding a second one that size is a hard sell.
So when evaluating every answer gets expensive, teams are left with three choices: check a sample of production traffic, check only before launch, or accept the extra bill.
A factory that inspects one car in five will ship the other four unchecked. Sampling is not a quality policy; it is a budget policy.
What a decision model does differently
A large language model is like an essay writer. Ask it to check an answer and it writes a paragraph of reasoning, then a verdict. You pay for every word.
A decision model is like the marker on a multiple-choice exam. You give it the material and a statement, and it returns one number: how likely the statement is to be true. No essay. And it can answer many statements about the same material in a single call.
- Source
- "Canberra is the capital city of Australia."
- AI answer
- "The capital of Australia is Sydney."
- Statement
- "The answer is supported by the source."
- Decision model
- 0.03 → almost certainly not supported. Flagged.
That is a much smaller job than writing an essay, which is where the cost difference comes from. By our estimate from the tokens it was billed for, a check from the decision model cost 50 to 100 times less than the same check from a frontier LLM judge.
Cheap is only useful if it is also right. So the real question was never "is it cheaper?" It was: is it good enough to trust, and what do you do on the cases where it isn't sure?
The idea we evaluated: a triage nurse
A hospital does not send every patient straight to the senior consultant. A triage nurse sees everyone first, handles the clear cases, and sends the uncertain ones up.
We set up evaluation the same way. The decision model sees every answer. Because it returns a probability, we know how sure it is. When it is confident, its verdict stands. When it is unsure, the answer goes to a strong LLM judge for a second opinion.
Check every answer. Escalate uncertainty.
- ConfidentKeep the first-line verdict.
- UncertainSend to a strong LLM judge for a second opinion.
Record the evaluation evidence
Coverage and escalation are different numbers.
- First evaluation
- 20% sampling example20 answers → LLM judge
- Adaptive evaluation100 answers → decision model
- Second opinion
- 20% sampling exampleNo escalation stage
- Adaptive evaluation13 uncertain answers → strong judge
- Answers evaluated
- 20% sampling example20 of 100
- Adaptive evaluation100 of 100
- Expensive judge calls
- 20% sampling example20
- Adaptive evaluation13
Numbers from our RAGTruth evaluation at the default dial. The punchline: 5× more coverage, with fewer expensive judge calls.
How sure is "sure enough"? That is a setting we call the confidence dial. Turn it up and more answers get a second opinion, which costs more but is more accurate. Turn it down and it costs less.
How we evaluated it
We used two public research datasets of AI answers, where humans had already marked which answers were made up (a "hallucination") and which were faithful to their source.
- RAGTruth
- 900 real answers written by GPT-4 and GPT-3.5, the hardest slice we could find. Their mistakes are rare (about 1 in 10) and subtle.
- HaluEval
- 1,500 question-answering and summary cases, half right and half made up, each pair built on the same source text.
- Same answers for all
- Within each dataset, every checker saw exactly the same answers. We compared the decision model, a small LLM judge, a strong LLM judge (GPT-4.1), and the triage setup.
Two traps nearly fooled us, and they are worth knowing about if you ever run an evaluation like this.
The order trap. One dataset is stored in a fixed pattern that repeats every six rows. Our first sample took every eighteenth row, which quietly picked answers from one model only. The results looked great, and they were wrong. We now shuffle every sample.
The ruler trap. In HaluEval, the correct answers are often one or two words ("Delhi") while the made-up ones are full sentences. Measuring the length of the answer, and nothing else, separated right from wrong almost perfectly (0.97 on a scale where 1.0 is perfect and 0.5 is guessing). So we kept only pairs where the right and wrong answers are about the same length, and checked that length now tells you nothing.
What we found
On its own, the decision model kept pace with a frontier judge
On the hardest slice, the decision model scored the same as GPT-4.1: 70.5% against 70.6% (balanced accuracy, where 50% is a coin flip). It did that at 50 to 100 times less per check. A smaller, cheaper LLM judge (GPT-4o mini) did clearly worse than both, whatever prompt we gave it.
Together, they beat either one alone
This was the surprise. The two models make different mistakes, so when the decision model passes its unsure cases to GPT-4.1, the pair is more accurate than GPT-4.1 checking everything: 73.6% and 74.3% in two separate runs, against 70.6%. That gap held up statistically both times.
And it was cheaper. The triage setup sent only 40% of answers to the expensive judge. With the dial turned lower (13% sent up), it matched GPT-4.1's accuracy for about one-eighth of the judge bill.
The second opinion has to come from a strong judge
We first tried the obvious budget option: the decision model plus a small, cheap LLM judge. It did not help. On the hard cases the decision model passed up, the small judge was barely better than a guess.
It makes sense once you say it out loud. The nurse only calls the consultant for the difficult cases. If the consultant is a trainee, calling them achieves nothing. Because the strong judge now only sees a small share of the answers, you can afford a good one.
It held on a second dataset
On HaluEval the result repeated. The decision model (80.1%) and GPT-4.1 (80.8%) were statistically level. At our default dial, the triage setup scored 80.4% while sending just 8% of answers to the judge: the same accuracy for about one-twelfth of the judge bill.
Accuracy, with the test conditions attached.
Balanced accuracy (%)
Jev 1.13.0 + GPT-4.1. RAGTruth: GPT-4/GPT-3.5 test answers. HaluEval: 375 QA pairs + 375 summary pairs, length-matched within 25%. Balanced accuracy averages sensitivity and specificity. Chance reference: 50%.
Split by task, the picture is more interesting. On short questions, GPT-4.1 was a little better (90.9% against 88.3%), and this is where the dial earns its place: turn it up to send 17% of answers to the judge, and the triage setup reaches 90.4%, level with GPT-4.1, for about one-sixth of the bill.
On long summaries it was the other way round: the decision model was slightly ahead (71.9% against 70.7%). And both of them found summaries hard, which brings us to the limits.
What it is not good at, yet
Subtle mistakes in long text. On summaries, where one wrong detail hides in a paragraph, both the decision model and GPT-4.1 scored only around 71%. Neither is a finished answer there, and a second opinion from the same kind of judge did not add much.
False alarms. On the hardest slice, where only 1 answer in 10 was wrong, it raised about three flags for every real problem it caught. That is fine for "look at this", not for "block this automatically".
It needs a clear question. It answers "is this statement true?", so the statement has to be written plainly and the right way round. "The answer is supported by the source" and "the answer contains unsupported claims" mean opposite things.
The verdict
Across the article’s 2,400 labelled answers, the default setting evaluated every answer at roughly one-tenth of the strong-judge bill. Reported balanced accuracy was 72.8% versus 70.6% on RAGTruth and 80.4% versus 80.8% on HaluEval. These results apply to the datasets and conditions described below.
Decision models are not a replacement for LLM judges. They are a first line that makes checking everything affordable, with a strong judge kept for the cases that need one.
Look at what share of your AI's production answers are checked today.
If the honest answer is "a sample" or "none", cost was probably the reason. That reason is going away.
What we built into AgentGuard: Adaptive Evaluation
A decision model on its own is an engine, not a system. What made these results possible was the process around it, and that is what we have built into AgentGuard. We call it Adaptive Evaluation.
Evaluate every answer, measure confidence, escalate uncertain cases to a strong judge, and record the evidence.
Jev is the first engine behind it. If a better decision model, a small language model or a classifier becomes the right first line tomorrow, it slots in, and everything around it stays the same.
- ⚡ Decision-ready evaluators
- 15 evaluators in the AgentGuard library now come with a plain yes/no question written for decision models. Connect a decision model and they switch on. Nothing to rebuild.
- A confidence dial
- Each evaluator has a setting for how sure the decision model must be. Unsure cases go to your chosen LLM judge automatically.
- Many checks, one call
- When several evaluators check the same conversation, AgentGuard asks all their questions in a single call, so adding checks barely adds cost.
- Never a blind spot
- If the decision model is ever unavailable, the evaluation falls back to the LLM judge. Nothing goes unchecked.
- Not tied to one vendor
- Jev is the first engine we support. The questions, the dial, the escalation and the evidence belong to AgentGuard, whichever model does the first check.
- Visible reasoning
- Every decision is recorded with the question asked and the answer given, so you can see why an answer was flagged.
Method and limits
Datasets. RAGTruth (MIT licence): all 900 test answers written by GPT-4 and GPT-3.5, 88 of them labelled as hallucinations by human annotators. HaluEval (MIT licence): 375 question-answering and 375 summarisation pairs, each with one correct and one hallucinated answer on the same source, kept only where the two answers are within 25% of each other in length. The question-answering sample covers almost every such pair in the dataset (375 of 384), so it is a particular subset rather than a random one.
Models. Decision model: TypeSafe Jev (jev-1.13.0). LLM judges: GPT-4o mini and GPT-4.1, asked to list every claim in the answer and check each against the source. The decision model was asked one plain statement: "The output contains information that is not supported by the context."
Scoring. Balanced accuracy (the average of the share of problems caught and the share of good answers passed), so that a model cannot score well by always saying "fine". Differences we call real held up in a paired statistical test (95% bootstrap interval). Costs are list prices for RAGTruth. The HaluEval judge ran at batch pricing, so we compare it as a share of the judge's own bill. The decision model's cost is estimated from its billed input tokens.
Limits. The decision model is not fully deterministic: we ran it twice on RAGTruth (2% of answers changed verdict, the conclusions did not) and once on HaluEval. Two datasets, one kind of check (is the answer supported by its source). Your traffic will have its own mix, so measure on it before you set the dial.