Blog

Eight Rules for Benchmarking a Model You're About to Trust

Engineering, Sep 29, 2026, 12 min read

Eight Rules for Benchmarking a Model You're About to Trust

We ran one request through a new model five times. Median: 8.084 seconds.

The next day we ran the identical request again: same payload, same model, same code path. Median: 0.669 seconds.

Both numbers are real. Neither is wrong. And if we had run that request once, like most model comparisons do, we would have published whichever one we happened to get.

Swapping the model behind a decision looks like a small change. A prompt, a client, a parsed response, an afternoon. What makes it not small is where these decisions sit: route a request to a fast lookup when it needed a full investigation and the user gets a confident, shallow answer. Decide an agent can service a task when it holds credentials for one cluster out of nine, and the failure surfaces much later, somewhere that looks unrelated.

These are the eight rules we put between a decision model and production. We built them while evaluating TypeSafe's Jev against our own LLM classifiers, and every one is here because skipping it would have let us believe something comfortable.

1. Freeze real requests, not written ones

Every benchmark starts from real traffic. For the capability checker that was 10 real Dev EU requests spanning 2026-09-15 to 2026-09-23: 15 checks, 36 tasks, 215 agent calls. Not requests we wrote to be representative. Requests that happened, carrying the context they actually carried.

Then freeze them. The exact capability text, datasource metadata and task batch from each call, replayed byte for byte on every run of every arm. If one side sees slightly different context the comparison is void, and nothing in the output will tell you.

Ten requests is a small sample, and we say so wherever a result is close. It is small because every row needs a human to decide what the right answer was, and ten reviewed rows beat two hundred unreviewed ones. Sample size is the honest reason we decline to call a near-tie.

2. Run everything three to five times

Five repeats per request per arm on the capability checker, three on the token benchmarks, 195 total valid runs on intent routing.

Repeats are not about averaging out noise. They are about whether a model is consistent, which for a routing decision matters nearly as much as being right.

Consistency across five identical runsBars. The current LLM checker changed 27 of 507 decisions across five identical runs, 5.3 percent. Jev per-agent changed 4 of 507, and Jev per-task 4 of 507, both 0.8 percent.Current LLM checker27 of 507 decisions (5.3%)Jev per-agent4 of 507 (0.8%)Jev per-task4 of 507 (0.8%)decisions that changed across five identical runs
Same frozen inputs, five times. The current checker runs at temperature 1.0, which is where most of this comes from.

Identical inputs, five times. Our checker runs at temperature 1 and changed 5.3% of its decisions; Jev changed 0.8%. One run per request cannot tell "right 80% of the time" apart from "right 80% of the time by flipping," and those are different systems with different failure modes.

Then there is the case from the top of this post:

One request, ten runs, three defensible answersLog-scale ranges. Old flow median 1.448 seconds. Jev’s first five runs median 8.084 seconds with four between 7.379 and 14.864. The recheck five median 0.669 seconds, range 0.613 to 1.163. Combined median of all ten, 1.072 seconds.0.5s1s2s5s10s15sOld flow5 timed runs1.448sJev v2, first five runsfour ran 7.379-14.864s8.084scause still unknownJev v2, recheck on 2026-09-24same payload, same model0.669sAll ten Jev v2 runswhat one number would report1.072srecorded duration, log scale
REQ 11b, unchanged payload and model. A single run could have reported 0.613s or 14.864s, and both would have been true.

Unchanged payload, unchanged model. The first five runs sat at a median of 8.084 seconds with four of them between 7.379 and 14.864. The recheck the next day sat at 0.669, slowest 1.163.

We publish the combined median of all ten, 1.072 seconds, and we publish the split beside it, because the cause of that first batch is still unknown. Picking either half would have been a story. Publishing both is the measurement.

3. Record every run, and leave the gaps as gaps

Each run saves its decisions, the route taken, token counts, billed cost, wall duration, and the provider's response ID.

The response ID is the field everyone skips and the one that makes the rest trustworthy. It means any number in any summary can be walked back to the exact provider call that produced it, months later, by someone who was not in the room. Without it you have a spreadsheet that is true because you say so.

Missing measurements stay blank. Not zero, not an estimate. Our intent-routing benchmark has 195 valid records but only 135 saved timings, because 60 earlier runs predate us capturing duration. So the timing charts are drawn from 135 points and say so, rather than quietly filling 60 holes with something plausible.

4. Publish the runs you threw away

Two exclusions from intent routing, both written into the result:

Four attempts at one request ran against the wrong Gemini version. Excluded, re-run. One Jev call returned HTTP 520; we retried it successfully and excluded the failed attempt from that request's five valid repeats.

Neither is interesting, which is the point. Exclusions are suspicious only when they are invisible, and a benchmark reporting zero discarded runs is usually one that was not watched closely. Writing down the boring ones is what makes it credible when you say there were no others.

5. Say what the numbers do not prove

This is the section I would most like other people's benchmarks to have.

For the capability checker, both flows ran from inside the same pod through the same gateway, on the same frozen inputs, inside an eleven-minute window. That is a fair latency comparison and we present it as one.

For intent routing it is not, and we label it everywhere it appears. Our old-flow timer covers classifier initialisation, three parallel calls and replay cleanup. The Jev measurement covers an HTTP POST and excludes key lookup. Some runs were measured in-pod, some from a local client. Those are different quantities of work.

We also did not test final answer quality anywhere in the intent-routing benchmark. We measured which route each configuration chose. Whether the answer that came back down that route was any good is a question we did not ask, and a reader who assumes otherwise has been misled by us, not by themselves.

If a benchmark has no limitations section, it has limitations.

6. Agreement with production is not correctness

The subtlest trap here, and the one that nearly cost us a working model.

When you replay a real request you know what your current system decided. Treating that as truth is easy to compute and feels objective. It measures the wrong thing. Your current system's decisions are observed behavior, not reviewed answers. Ours had been handing GKE tasks to the AWS agent for months. Score agreement and the ceiling for any new model is reproducing your mistakes exactly, while every case where it is right and you were wrong counts against it.

So agreement is a pointer to where to look, never a score. For a score you need a key.

Ours took 4 read-only audits of each agent's real tools, deployed prompts, live datasource configuration and the benchmark traces. 134 (task, agent) pairs labelled CAN, PARTIAL, CANNOT or UNSURE, each citing its evidence. No LLM judged any label: a model grading a model's routing decision inherits exactly the ambiguity you are trying to measure.

Two things only an audited key can show you. Two of our tasks have no capable agent at all, so the correct answer is to assign nobody, something agreement scoring cannot even represent, because production assigned somebody. And PARTIAL had to exist: an agent that keeps two hours of events when the task wants seven days is neither a yes nor a no, and forcing it into a binary scores a reasonable refusal as an error.

Here is what trusting agreement would have cost us. Jev's first intent-routing configuration agreed with production on 27 of 50 runs. As a score, that is a rejection. As a pointer, it said go look at requests 4, 5, 7 and 10: all counts and lists, all answering from context instead of researching, in every single run. The cause was our own routing rule, which never said a complete list requires research. We rewrote the rule. Agreement went to 85 of 85.

We would have rejected a model that worked and kept the defect that caused the disagreement.

7. Trace every error to a cause

The rule that changed the conclusion of the whole evaluation.

Counting errors is not enough. Every wrong decision gets traced to why, and the categories have to include ones that do not blame the model, or you will never find out that the model was not the problem.

What caused each wrong assignmentGrouped bars for four causes across four routers. Descriptions that claim an ability the agent lacks: 11, 10.0, 7.0 and 8.6 per run. Descriptions missing scope or a cannot-do line: 8, 5.6, 8.0, 8.0. A render step leaked from the planner: 0, 0.6, 2.8, 4.8. Router error where the description was clear: 0, 0.4, 0, 0.Recorded productionOld per-agent LLMsJev per-agentJev per-taskThe description claims an ability the agent lacks11.010.07.08.6The description lacks scope, or a "cannot do" line8.05.68.08.0A render step leaked from the planner00.62.84.8Router error, where the description was clear00.400wrong assignments per run, mean of 5 runs
Every wrong assignment in round one, traced to its cause. The bottom row is the only one that blames the router, and it is empty for three of the four flows.

Three routers sat within five points of each other around 70% precision. Counted, that reads as "the new model is no better than what we have." Attributed, it reads completely differently: router error, where the description was clear, is 0.4 per run at worst and zero for three of the four flows. Everything else traces to capability descriptions claiming abilities the agents lacked, omitting which cluster they covered, or to a render step that leaked out of our planner.

We fixed the descriptions and re-ran. Our own incumbent checker gained 14.5 points of precision from a prompt-only change. Without the attribution step we would have shipped a post comparing two models and never found the actual bottleneck.

Attribution drags up things outside the benchmark, too. Classifying 507 decisions means reading code, and we found a formatter appending a colon to empty format specs (so every literal {…} in every prompt reached the model malformed) and a cache keyed on datasource UID but not instance, quietly sharing one catalog between two environments. Neither is a benchmark result. Neither would have been found by counting.

8. Report the curve, not a point

Precision and coverage as the threshold movesLine chart over thresholds 0.4, 0.5 and 0.6. Jev per-agent precision rises from 91 to 94.8 to 97.3 percent while coverage falls from 91 to 85.3 to 84.1. Jev per-task precision rises from 93.7 to 97.2 while coverage falls from 88.2 to 84.7.80%85%90%95%100%0.40.50.6threshold: the score above which Jev claims the agent can do the taskJev per-agent: precisionJev per-agent: coverageJev per-task: precisionJev per-task: coverage91%94.8%97.3%91%85.3%84.1%93.7%97.2%88.2%84.7%
A scored model does not have one accuracy. Every point here is the same model on the same requests. Only the cutoff moved.

A scored model has no single accuracy. At threshold 0.4 it is 91% precise covering 91% of tasks; at 0.6, 97% precise covering 84%. Same model, same requests, same run. Only the cutoff moved.

Report one number and you have hidden a product decision inside it. Report the curve and the decision stays where it belongs: take coverage for a reversible suggestion a person reviews, take precision for a decision feeding something expensive and let it abstain.

One caveat we hold ourselves to: both of those thresholds were chosen on the same ten requests they are measured on. They are fitted until a fresh sample confirms them, and that belongs in the result, not a footnote.

The whole harness

  1. Real requests from traces, frozen with their context, replayed byte for byte.
  2. Three to five repeats per request per arm. Consistency is a result, not noise.
  3. Decisions, tokens, cost, duration and response ID per run. Gaps stay gaps.
  4. Exclusions published, including the boring ones.
  5. An explicit statement of what the numbers do not prove.
  6. Agreement as a pointer, never a score. Build an audited key, allow PARTIAL, let no model grade it.
  7. Every error traced to a cause, with categories that can exonerate the model.
  8. The precision-coverage curve, not a point, and a note when the thresholds were fitted.

None of this is sophisticated, which is deliberate: it is cheap enough to run for a decision as small as "which agent can service this task." Two of the three most valuable things it has told us were not about models at all. One was that sixteen text files were wrong. The other was that a rule we wrote ourselves was vague.

That is the argument for a harness. Not that it ranks vendors. That it tells you when you are about to buy a model to fix a problem a text file would have fixed.


The evaluation this came out of, with the full measurements and every caveat, is in We Benchmarked Jev Against Our Own Classifiers.

All posts