Blog

We Benchmarked Jev Against Our Own Classifiers. Every Router Stalled at 70%.

Engineering, Oct 1, 2026, 17 min read

We Benchmarked Jev Against Our Own Classifiers. Every Router Stalled at 70%.

Three routers. Two vendors. Architectures with nothing in common. All of them wrong about a third of the time, and wrong in the same places.

That is the result that turned a routine vendor evaluation into the most useful week of engineering we have had this quarter. When three models that share no code and no training fail identically, you are no longer measuring models.

Here is the whole thing: what we tested, what we measured, where our own conclusion turned out to be wrong, and the fourteen points of precision our own text files had been quietly costing us.

The decisions nobody writes blog posts about

Dalton makes a lot of small decisions before it makes any interesting ones.

Is this request a quick lookup or a full investigation? Which of our agents can actually service this task? Does it pass our quality rules? Should we offer an investigation at all?

None of that is the product. It is the plumbing in front of the product, and all of it ran on general-purpose LLM calls: a prompt, a schema, a model, a parsed answer. It works. It is also slow and expensive in a way that compounds, because every one of those calls sits on the critical path of a person waiting.

So when TypeSafe shipped Jev, a model built for narrow scored decisions instead of open-ended generation, we pointed it at five of them.

Decision What it answers How it runs today
Intent routing Fast lookup or full research? Which sources? 3 parallel LLM calls
Batch capability checker Which agent can service this task? 1 LLM call per agent, ~16 per check
QC rule validation Does this request pass our quality rules? 1 LLM call
QC mode validation Is the requested mode valid? 1 LLM call
Investigation-offer detection Should we offer an investigation? 1 LLM call

Most of this post lives in the capability checker, because that is where it got interesting.

Jev, in one request

Jev is not a chat model, and that difference drives everything below.

You send a state (one shared context) and a set of questions about it. A yes/no question is a noul. What comes back is not a verdict. It is a probability.

Here is the shape of one real request, asking whether our kubernetes-edge agent can handle four tasks from "Are there any current issues in the demo-system namespace?" Its own description is the state, each task is a question, and what comes back is four numbers.

One call: a shared state, four questions, four probabilitiesOne request diagrammed. The agent description is the shared state feeding four task questions. The answers are 0.23, 0.03, 0.02 and 0.02, all below the dashed 0.5 threshold; the 0.23 is marked as one that should have been a yes.statekubernetes-edge4 connected clusters3 things it can do7 it cannotone question per taskprobability of yesthreshold 0.5t1Which cluster holds it?0.23should be yest2List unhealthy pods0.03t3Get warning events0.02t4Query restart metrics0.02
The whole request for one agent against one four-task check: 1,282 input tokens, 72 out, $0.0000538. Three of those answers are right and emphatic. The 0.23 is a miss: our key says this agent can check which clusters hold a namespace.

The 0.03 on t2 is right and emphatic: listing pods is in that agent's cannot list. The 0.23 on t1 is a miss: our answer key says this agent can check which clusters hold a namespace.

Two consequences, both of which come back later. The output is a number, so its size is fixed: across all twelve intent-routing requests Jev returned exactly 55 output tokens every time, where our LLM flow returned 116 to 150. And because the answer is a probability, where you cut it is your decision, not the model's. That turns out to be worth six points of precision.

We ran Jev in two layouts. Per-agent mirrors today: one request per agent, each task a question. Per-task flips it: one request per task, each of the sixteen agents a question. A check with 4 tasks and 16 agents asks 64 questions either way: per-agent sends 16 requests of 4 questions, per-task sends 4 of 17.

The dataset is 10 real Dev EU requests from 2026-09-15 to 2026-09-23: 15 checks, 36 tasks, 215 agent calls, captured from traces and replayed frozen. Round one was 150 runs (10 requests × 3 flows × 5 repeats) on 2026-09-23 between 13:36 and 13:47 UTC, from inside the Dev EU snakur-2 pod so both flows took the production network path. Every call returned HTTP 200.

The easy part

Routing latency, three flowsLog-scale bars. Routing time p50: old flow 8.3 seconds, Jev per-agent 0.77, Jev per-task 0.56. p90: 23.5, 1.59, 1.34. Time to first claim p50: 4.47, 0.46, 0.44. One call p50: 4.8, 0.48, 0.47. Slowest call: 19.1, 1.2, 0.81.Old per-agent LLMsJev per-agentJev per-task0.5s1s2s5s10s25sRouting time, p508.3s0.77s0.56sRouting time, p9023.5s1.59s1.34sTime to first claim, p50when the first task can start4.47s0.46s0.44sOne call, p504.8s0.48s0.47sOne call, slowest seen19.1s1.2s0.81sseconds, log scale, 150 runs from inside the Dev EU pod, 2026-09-23
The old flow fans out one model call per agent (21.5 calls for an average request) and waits for all of them. Jev per-task asks the same questions in 3.6 requests.

Routing takes 8.3 seconds today at p50. Jev per-task takes 0.56. The number I care about more is time to first claim, how long before the first task can start working, which goes from 4.47 seconds to 0.44.

Look at p90 if you want to know how this feels to a user. Today it is 23.5 seconds, and the slowest single call we recorded took 19.1. Jev's slowest call in the entire run was 1.2 seconds.

Most of that gap is not model speed. It is fan-out. The old flow makes 21.5 calls for an average request and cannot finish until the slowest one returns. Per-task makes 3.6.

Cost per decision, four decision pointsLog-scale paired bars for four decisions. Intent routing: $0.00147 against $0.0000333, 44 times cheaper. QC rule validation: $0.00149 against $0.000217, 6.9 times. QC mode validation: $0.00148 against $0.000181, 8.1 times. Investigation-offer detection: $0.00303 against $0.0000818, 37 times.Current LLM classifierJev$0.0001$0.0005$0.001$0.005Intent routingENG-4797 · 43× cheaper$0.00148$0.000034QC rule validationENG-4798 · 6.9× cheaper$0.001491$0.000217QC mode validationENG-4799 · 8.1× cheaper$0.001476$0.000181Investigation-offer detectionENG-4801 · 37.0× cheaper$0.003028$0.000082billed USD per decision, log scale
Billed cost, not estimated. The ratio varies from 6.9× to 44× because the saving is mostly price per token, not tokens. See the paragraph below.

Cost falls everywhere, but the ratio swings from 6.9× to 43×, and the reason is worth knowing. On QC rule validation Jev used only 1.09× fewer tokens (nearly identical work) and still cost 6.9× less. That gap is price per token, not efficiency. On investigation-offer detection both effects compound: 5.4× fewer tokens (327,263 against 60,390) and the cheaper rate, for 37×.

Intent routing shows the fan-out saving directly in tokens. Three parallel classifier calls burn 4,697-6,419 input tokens per decision. One Jev call answering the same three questions burns 677-1,247.

Consistency across five identical runsBars. The current LLM checker changed 27 of 507 decisions across five identical runs, 5.3 percent. Jev per-agent changed 4 of 507, and Jev per-task 4 of 507, both 0.8 percent.Current LLM checker27 of 507 decisions (5.3%)Jev per-agent4 of 507 (0.8%)Jev per-task4 of 507 (0.8%)decisions that changed across five identical runs
Same frozen inputs, five times. The current checker runs at temperature 1.0, which is where most of this comes from.

Consistency is the result I expected to ignore and now think is the most underrated number here. Same frozen inputs, five times: our checker changed 5.3% of its decisions, Jev changed 0.8%. We run at temperature 1, which explains most of it. But a router that is right 80% of the time stably is a system you can reason about. A router that is right 80% of the time by flipping is a different system, and a worse one.

So far this is a clean win. Here is where it stopped being one.

The wall

Agreement with production tells you how similar two routers are. It does not tell you which is right, and ours was wrong in places, including handing GKE tasks to the AWS agent. So we built an answer key.

4 read-only audits went through each agent's real tools, its deployed prompts, the dev org's live datasource configuration, and the benchmark traces. 134 (task, agent) pairs labelled CAN, PARTIAL, CANNOT or UNSURE, every label citing its evidence. No LLM judged any label. Two tasks turned out to have no capable agent at all (one needs a cluster no Grafana covers, the other is an internal render step), so the right answer for both is to assign nobody.

Scored against that key, Jev per-agent came in at 69.6% precision.

We were ready to write that up as a loss. Then we scored our own LLM checker on the same key: 74.1%. And production itself: 70.8%.

Three routers within five points of each other, all wrong about a third of the time. At that point the interesting question is not which model is better. It is what all three have in common.

We stopped comparing models and traced 507 decisions

So we did the unglamorous thing. Every wrong assignment, traced to a cause.

What caused each wrong assignmentGrouped bars for four causes across four routers. Descriptions that claim an ability the agent lacks: 11, 10.0, 7.0 and 8.6 per run. Descriptions missing scope or a cannot-do line: 8, 5.6, 8.0, 8.0. A render step leaked from the planner: 0, 0.6, 2.8, 4.8. Router error where the description was clear: 0, 0.4, 0, 0.Recorded productionOld per-agent LLMsJev per-agentJev per-taskThe description claims an ability the agent lacks11.010.07.08.6The description lacks scope, or a "cannot do" line8.05.68.08.0A render step leaked from the planner00.62.84.8Router error, where the description was clear00.400wrong assignments per run, mean of 5 runs
Every wrong assignment in round one, traced to its cause. The bottom row is the only one that blames the router, and it is empty for three of the four flows.

Read the bottom row first. Router error, where the description was clear: 0, 0.4, 0, 0.

Out of roughly seventeen wrong assignments per run, the number you can blame on a model misreading something it could have read correctly rounds to zero. For three of the four flows it is zero.

Everything else is ours. Descriptions claiming abilities the agents do not have, seven to eleven times per run. Descriptions that never say which cluster an agent covers or what it cannot do, five to eight. A render step that leaked out of the planner and got offered to all sixteen agents.

Sixteen text files

A router sees exactly two things: the agent's USAGE.txt, and a metadata header with the datasource name, type and UID. So we read all 16 of them against what the agents can actually do.

1 was accurate. Every other one claimed at least one ability its agent does not have, and most never said what their agent cannot do at all.

The causes were structural, which is why nobody had noticed:

  • No text carries the org's configuration. The AWS and GCP domain agents load USAGE.txt as a fixed file (aws/base_domain_agent.py:88, gcp/base_domain_agent.py:94), and the metadata header adds only name, type and UID (utils/datasources/common.py:139-177). So no router can learn that this AWS account has no EKS cluster, that dalton-dev exports no workload logs, that the prod clusters sit outside the GCP datasource's scope, or that both Loki instances are empty. Our older single AWS agent had already solved this: it renders real account IDs into its text (aws/agent.py:184-201). Nobody carried the pattern forward.
  • Some "cannot do" lines exist in the wrong place. For kubernetes-edge, memgraph-cypher and mongodb-atlas they live only in planner cards. The capability checker never sees planner cards. Neither does Jev.
  • A render step leaked into routing. "Render change timeline" left the planner without action=render_section (plan_generator.py:526-532), so it went to all sixteen agents. Our own rule that other tasks' outputs count as already available (CAN_SOLVE_BATCH.txt:9-13) makes combining results look easy to any agent whose text mentions analysis or timelines.
  • Nothing documents the 2-hour window. GCP logs, audit entries and metrics, CloudWatch and Atlas metrics each cover at most two hours per call. "The last 48 hours" is therefore 24 to 84 calls. No text says so, so no router can price it.

Tracing those also turned up two bugs that have nothing to do with routing. safe_format.py:132-133 appends a colon even when the format spec is empty, so every literal {…} in a prompt reaches the model as {…:}. Every LogQL example we ship was arriving malformed. And a Loki label cache keyed on datasource UID but not instance (discover_loki_label_values.py:280) had two environments sharing one label catalog, because both use the UID loki.

Neither is a benchmark result. We found them because classifying 507 decisions makes you read code you would otherwise trust.

What accurate descriptions are worth

Round two ran the next morning, 08:16 to 08:27 UTC, same pod, same ten requests. Two things changed, and we separated them deliberately so we would know which one mattered.

First, request shape: TypeSafe documents a structured form we had not used. Changing only that moved precision from about 70% to 72%. Essentially nothing, which is exactly what made it worth running, because it ruled out formatting as the cause.

Second, corrected descriptions: short, accurate statements of what each agent can and cannot do, and which clusters it covers.

Precision against coverage, before and after the descriptions were correctedScatter plot. Four points sit between 69.6 and 74.1 percent precision, all using today’s capability text. Five points sit between 88.6 and 97.3 percent, all using corrected descriptions. A dashed line at 75 percent separates the two groups.65%70%75%80%85%90%95%100%80%85%90%95%100%the ceiling every router hit on today’s descriptionsproduction, as recordedLLM checker, today’s textJev, prompt-styleJev, documented shapeLLM checker + factsJev per-agent + facts @0.4@0.5@0.6Jev per-task + facts @0.5@0.6Coverage: share of tasks given a capable agent →Precision ↑
The same ten requests, scored against the same answer key. Below the dashed line, every router is reading today’s capability text. Above it, every router is reading corrected text. The gap between the two clusters is the finding.

Everything below the dashed line is a router reading today's text. Everything above it is the same routers reading corrected text. Nothing else changed.

Our own LLM checker gained 14.5 points of precision from a prompt-only change. Nobody touched the model, the temperature or the request shape. Jev's wrong assignments fell from 17.8 to between 1.0 and 3.0 per run. Both started correctly leaving the two ownerless tasks alone, which neither had reliably done before.

If you take one number from this post, take that one. The incumbent system we were trying to replace got most of the way to the challenger's accuracy by having its input files corrected.

Where you put the cutoff is a product decision

Precision and coverage as the threshold movesLine chart over thresholds 0.4, 0.5 and 0.6. Jev per-agent precision rises from 91 to 94.8 to 97.3 percent while coverage falls from 91 to 85.3 to 84.1. Jev per-task precision rises from 93.7 to 97.2 while coverage falls from 88.2 to 84.7.80%85%90%95%100%0.40.50.6threshold: the score above which Jev claims the agent can do the taskJev per-agent: precisionJev per-agent: coverageJev per-task: precisionJev per-task: coverage91%94.8%97.3%91%85.3%84.1%93.7%97.2%88.2%84.7%
A scored model does not have one accuracy. Every point here is the same model on the same requests. Only the cutoff moved.

Given equal descriptions, Jev matches or beats our checker depending entirely on the threshold. At 0.4 it is 91% precise covering 91% of tasks, against 89% and 91% for the LLM checker, while running 9× faster per-agent, 14× faster per-task, and 20× cheaper in both. At 0.6 it makes about one wrong assignment per run, but covers 84%.

Those are not three accuracies. They are three products, and choosing between them is our job: for a reversible suggestion a person reviews, take the coverage. For a decision feeding something expensive, take the precision and let it abstain.

What it declines is informative too. Four of the five uncovered tasks have only partial owners (seven days of node events where memgraph-cypher keeps about two hours, past-window logs where kubernetes-edge returns only the latest lines), and Jev says no to a partial fit, which I think is defensible. The fifth is our wording: our corrected text says kubernetes-edge cannot search logs, Jev read "gather error logs" as a search, and its score fell from 0.9 to 0.2. One ambiguous clause, six points of coverage.

We also built a comparison question to choose between our two Grafana instances. It was dead weight: with each instance's cluster written into its description, the plain yes/no answers already rejected the wrong one, and the comparison never overruled a claim in 61 chances. The one instance error that survives cannot be fixed at routing time at all, because the cluster it needs is only known from a previous task's result.

Then it happened again

While that was running, the intent-routing benchmark produced the same finding independently.

Jev's first configuration agreed with production on 27 of 50 runs. Our instinct, again, was that Jev was wrong.

Which route each configuration chose, per requestA matrix of twelve requests by three configurations. Production chose research on nine requests and instant on three. Jev v1 chose answered-from-context on requests 4, 5, 7 and 10, and mixed on request 3 - every one of them a count or list request. Jev v2 matched production on all twelve.ResearchAnswered from contextInstantMixed across runs1234567891011a11bProductionthe route we take todayRRRRRRRIIRRIJev v1our first routing ruleRRMCCRCIIC--27/50Jev v2the rule, rewrittenRRRRRRRIIRRI85/85a count or list request · R = research · C = answered from context · I = instant · M = mixed across runs
Jev v1 did not fail randomly. On every count or list request it answered from the context it already had instead of going to look - because our rule never said those need research. The lit cells are the disagreements.

It was not failing randomly. Every count and every list request came back CONTEXT (answer from what you already have) where production went and looked. Our routing rule never said that a complete list or a count requires research. Jev followed the rule we wrote.

We rewrote it to name those cases and changed nothing else. 85 of 85.

Twice in one evaluation, in two unrelated benchmarks, what presented as a model limitation was something we had not written down clearly enough.

What we are not claiming

The corrected descriptions came from the same audit as the answer key. Same four audits wrote both. So those arms show what accurate descriptions can achieve, not a held-out test. Re-deriving the descriptions independently would be a stronger result and we have not done it.

Both thresholds were chosen on the same ten requests they are scored on. 0.4 and 0.6 are fitted until a fresh sample says otherwise.

The answer key is a proposal under review. 134 labels, several of them genuine judgment calls. Every precision number here moves if a reviewer disagrees.

Three of the five benchmarks report cost and tokens only. QC rule validation, QC mode validation and investigation-offer detection have complete billing evidence but their expected decisions are still under review, so no accuracy figure in this post comes from them. When that review lands the numbers may be unflattering. We will publish them.

Jev is not a drop-in replacement. On investigation-offer detection, the five cases where the answer is "yes, offer one" also require writing the offer's title and context. Jev scores the decision; it cannot write that text. Something else has to, and that cost is not in the 37× figure.

The latency numbers are recorded, not controlled. For the capability checker both flows ran from the same pod through the same gateway, which is fair. For intent routing they did not: our old timer covers classifier initialisation, three parallel calls and replay cleanup, while the Jev measurement covers an HTTP POST. Different quantities of work. We publish them as observed because a 1.249-2.935s band against sub-second responses is large enough to act on and not precise enough to quote.

What I would do first

We are moving ahead with Jev where it fits. It is faster, materially cheaper, far steadier run to run, and the threshold dial is a real advantage for decisions where abstaining beats guessing. The order is written down: fix the descriptions, re-run the same ten requests, then pick a threshold and send the uncertain band to the existing checker.

But the change with the biggest payoff in the entire evaluation was editing sixteen text files. It beat every model choice on the table, it improved the system we were trying to replace by 14.5 points on the way past, and it would have been the right move whichever vendor we picked.

We have argued before that your AI strategy is a data strategy wearing a costume. This is that argument with 507 traced decisions behind it.

If you are about to benchmark models for a routing or classification decision, spend the first day reading the metadata your models will read. Your ceiling may have nothing to do with the models.


The harness this came out of (frozen requests, published exclusions, and the rule that stopped us trusting our own baseline) is in How We Benchmark a Decision Before It Touches Production.

All posts