Verifier replay bench · frozen 20-case suite

Drilling for the cheapest
sufficient judge

A closure-contract verifier answers one question: does this evidence support closing this clause? Getting it wrong in the permissive direction promotes unsupported state. We ran the identical frozen suite from a 230M model on a 2017 GPU up to a 2.81T hosted frontier model, to find where sufficiency actually begins.

01 · The instrument

Five defects in the measuring stand, none found by the stand

Three were found by the operator noticing something that looked wrong; two were self-inflicted mid-run. Several produced no error at all and broke nothing visible, which is what made them expensive.

R2 · measurement floor

Two models tied to the second

Both local models reported exactly 58s. The client addressed the provider as localhost, which resolves to ::1 first; the refused connection cost ~2.05s per request. Constant across a 24× change in prompt size, so not prefill. 79% of the published latency was socket floor. Verdicts were unaffected — which is what made it invisible.

R3 · parser

A verdict invented from mid-thought

The parser took the last status token anywhere in the reply. A reasoning model whose thinking never terminated, musing “I think it is SATISFIED”, parsed as SATISFIED — a closure manufactured from a model that never issued one. Note the direction: it could only ever fabricate closure, never refuse one, matching the failure it was built to hunt.

R4 · noise floor

Temperature 0 is not deterministic in the cloud

Two identical runs disagreed on 12 of 20 cases for one model and 0–1 for most others. The local control reproduced exactly (6, 6, 6 — zero flips), so the variance is the provider, not the harness. Any cloud difference of one to two cases is inside the noise.

R5 · orchestration

A proxy for “the previous run finished”

A follow-on job waited for the results file to stop growing instead of asking whether the process was alive. A model taking 38s per verdict stalled the file past the threshold, so a second run launched onto the same GPU — each unloading the other’s resident model. 17 records were destroyed. Seven came back as HTTP 400 and read as “model incompatible”; retested individually afterwards, all of them loaded and answered fine. Same shape as R2: a proxy standing in for the thing itself.

R6 · scoring

A rate limit recorded as a wrong answer

Three calls returned HTTP 429. The scorer treated a row with no answer as UNPARSEABLE and charged 20 loss points, so two models were published at loss 20 whose real loss was 0 — and it manufactured a tidy false narrative that the newer generation was regressing. The scorer now excludes rows the model never answered, and reports how many cases the metrics rest on. Unevaluated is not failed.

Three of the four could only ever bias results in a direction that flattered the run — a latency floor that hid the cheap models, a parser that manufactured closure, and a corruption that looked like model incapacity. An instrument whose failures all point one way cannot be trusted to measure that direction, which is why they are reported here rather than fixed quietly.

02 · Capacity

There is no parameter count at which safety switches on

Each point is one local model. Vertical position is unsupported closure — how many times it closed a clause the evidence did not support, the one error the architecture exists to prevent.

A 1.7B reaches zero. A 7B posts six. That same 7B’s reasoning-tuned sibling posts zero. Scatter within a size band exceeds the trend across bands.
The question “what is the minimum stochastic mass for this operation” has no answer, because mass is not the governing variable.
03 · Reasoning

Thinking is a repair, not an upgrade

The same cloud models, run twice, with provider-side reasoning as the only variable. Lines falling to the left are models rescued by reasoning.

Reasoning rescues models that were failing and does nothing — or mildly hurts — for models already at ceiling. Where it hurts, it hurts by over-abstention, which is the cheap error, not the expensive one.

One instance was withdrawn. An apparent 207 → 6 repair turned out to be a bad draw: that model scores 7, 3, 207 across repeated identical runs, and its thinking-on score of 6 is simply its normal score. The effect survives where the baseline is stable — one model scored 15 three times running before dropping to 0 with reasoning, and the local olmo-3-7b instruct/think pair is deterministic.

04 · Scale

Two hundred times the parameters buys nothing

Cloud models, reasoning off, ordered by loss. The cheapest model tested ties the best score in the pool; two of the largest score worse.

Differences of one to two cases sit inside the measured noise floor, so the frontier band is not rank-ordered by this data. What is established is the absence of benefit, not the presence of harm: nothing in the pool beats a 20B, and a local 12B on a 2017 GPU matches the best hosted result exactly. The largest model tested carries 2.81T parameters by its published spec — roughly 140× that 20B and 234× the local 12B — and scores the same.

05 · The frontier

What passes the safety gate

Gate: zero unsupported closure and 4/4 paired discrimination. The paired test uses adversarial cases and their defect-removed twins, so a model that pattern-matches suspicion fails it, and so does one that is merely timid.

And what failure looks like

The table above is a list of passes, so on its own it hides the thing this contract exists to catch. These are the runs that closed a clause the evidence did not support — the error that promotes unsupported state, and the reason the loss function prices it a hundred times higher than being wrong the other way.

One clause pulls a false closure out of one run in five

Counted over every complete run in the study, on the cases where closing is never the admissible answer. Runs, not distinct models: the pool holds sibling checkpoints, quantisation variants and repeat configurations of the same line, so these are counts of measured artifacts. If unsupported closure were scattered model-specific error, the rates would be roughly level across cases. They are not.

The worst case is not the most complicated one. It is the one that offers a true fact and invites a larger conclusion: a fixture passed, therefore the general claim holds. It drew a false closure from , more than any other case in the suite. One passing instance is not an established general property.

Three unrelated models from three different families make exactly this jump, which is what separates a shared attractor from scattered noise. It also settles a tempting shortcut: one of them separates all four twin pairs and still closes an unsupported clause here. Telling a defect from its repaired twin and refusing to over-claim are different abilities, and a gate that tests only the first would have passed it.

Local and cloud are never ranked together — a 4B on this GPU and a hosted trillion-parameter model do not share a cost model.
06 · Stability

The most dangerous model is the intermittent one

Three identical draws per model. The local control is bit-stable; one cloud model is bimodal.
A verifier that is safe in four draws of five, and posts two false closures in the fifth, is worse than one that is reliably mediocre — it passes acceptance testing and then closes unsupported state in production.

Counting the flips was the wrong measurement

Several cases admit more than one correct answer, so a verdict can change between identical requests and cost nothing at all. Counted that way, two models below look similarly noisy. Split by whether the change leaves the admissible set, they are not alike in the slightest.

One model changes its answer on two cases and never leaves the admissible set: for a system that only cares which transitions are allowed, it behaves exactly like a deterministic one. Another changes on six, leaves the set every time, and twice lands on closing a clause the evidence does not support.

The distinction is the difference between a judge who words the same refusal differently and a judge who refuses on Monday and permits on Tuesday, given the identical file. Only the second is a hazard, and the raw flip count cannot tell them apart.

What counts as harmless here is a property of this contract, not of the words. Two of the five statuses are treated as interchangeable because under these rules both refuse to close, so either one leaves the state exactly where it was. A system that gave them different consequences — one asking for more evidence, the other ending the attempt — would have to redraw the line, and some of the drift counted as free above would stop being free. The unit to measure is not whether the answer changed, but whether what the answer authorises changed.

07 · Build identity

The same weights at a different quantisation are a different verifier

Local models reproduce exactly here — the control returned 6, 6, 6 with zero flipped cases — so a difference between two builds of one model is a difference, not a draw.

Qualification attaches to the artifact, not to the model name. A verifier accepted at one quantisation has not been accepted at another.

It is not a “more bits is better” law, which is what makes it awkward: two families got worse with a larger quantisation. The effect is real, model-specific and non-monotonic, so it cannot be predicted — only measured.

08 · Third arm

The paid route buys no accuracy over a local 12B

A separate arm, run through a hosted CLI verifier harness. It is not on the same axis as the other two: the harness adds its own read-only-verifier framing and returns a structured object instead of a bare status word. Provenance was confirmed on every call — the route reports which model actually answered, and reports nothing when none did.

The heavier tier scored worse than the lighter one, and every one of its errors ran the same way: it refused to close cases the oracle calls plainly satisfied — including the control case that exists to catch exactly that. Safe, but a verifier that can never say yes escalates everything.

One model in this arm was also scored on the cloud arm, as a check on the harness itself. It moved by a single case — which is precisely the measured noise floor, so the harness effect is not established and is indistinguishable from a draw at one run.

09 · Silent failure

Valid syntax did not imply stable judgement

One checkpoint, four ways of asking for the same verdict, three repeated draws each. Same provider, same credential, same frozen evidence, same scorer. The only thing that varies is the shape the answer must arrive in.

Asking for one bare word with no explanation scored a perfect zero three times over. Requiring a machine-readable format instead produced replies that were almost always well-formed and frequently different from each other.

This kills two comfortable explanations at once. It is not that the nested shape overloads the model — the minimal one-field version produced the single worst draw in the experiment. And it is not that the model needs room to think out loud before committing — the arm with no room at all, one word and nothing else, was the best of the four.

What the constrained arms lost was reproducibility, not one identifiable ability. But the raw count of unstable cases hides the distinction that decides whether instability costs anything. Several cases admit more than one correct answer, so a verdict can move between draws and stay entirely within the admissible set. That kind of drift is free. The kind that matters moves across the boundary — and the two behave nothing alike here.

Every disagreement in the one-word arm stayed inside the admissible set. Counting each structured arm separately, eleven of their fourteen drifting slots crossed out of it — seven distinct cases, since the two arms largely destabilise the same ones — several landing on the answer that closes a clause the evidence does not support.

So the one-word arm is not bit-stable, and this does not claim it is: it answered two cases differently across draws. It simply never used that freedom to leave the region where any answer is acceptable. The structured arms did, repeatedly. And they did it in much the same places: six of the eight cases that drift do so under both of them. All eight were already weak surfaces by an independent measure — every one of them had drawn false closures elsewhere in that census, at rates from to , and not one was a case the census found clean. Runs, again, not distinct models: that pool holds sibling checkpoints and quantisation variants of the same lines. So within this frozen suite the imposed format loosened surfaces that were already measurably weak rather than opening a new kind of hole. Outside these twenty cases it could do either, and nothing here speaks to that.

Across the structured arms, exactly one reply was malformed — well-formed. A monitor watching format compliance would have reported near-perfect health while the verdicts underneath it moved across the boundary that matters.

That is worse than malformed output, which at least announces itself. Here the contract the machine checks was satisfied completely and the contract that matters was not. Note what is and is not established: the split in this experiment falls exactly on whether a format was imposed. What happens behind that switch — constrained decoding, a separate provider path, a prompt transformation, something else — was not observed and is not claimed. Three of the four models tried elsewhere showed no such effect at all, so this is one checkpoint's behaviour, not a law.

The engineering reading survives that caution intact. The imposed format was trying to enforce, during stochastic generation, a property the runtime can check deterministically once generation is over. A reply either is one of five permitted words or it is not, and answering that question is one comparison. Doing it afterwards does not make the model more likely to comply — it makes non-compliance cheap, visible, and unable to reach the state it would have changed.

10 · One vendor, ten generations

The cheapest tier had already reached the judgement ceiling two generations ago

Ten fixed model IDs from one provider — no -latest aliases, because qualification binds to a configuration and an alias can be repointed at another checkpoint. Every record carries the provider's own reported version, and all ten matched the ID requested.

Every model in the ladder has zero unsupported closure. The differences between them are speed, tail risk, and how many thinking tokens they spend to reach the same answer.

The cheapest tier posts a perfect score with zero thinking tokens, while the most expensive tier spends tens of thousands of them, takes roughly 28× as long, and scores worse. Stated precisely: this is not a leaderboard win. On raw median latency the cheapest tier is actually beaten by 16 ms — by a later model with a worse score. The claim that survives a literal reading is about the ceiling: it was reached two generations ago, and nothing since has exceeded it. The only genuine regression is inside the cheap line itself, where a newer Lite is worse than the one it replaced.

The reasoning budget is genuinely getting cheaper. Across the three newest models that all score identically — every case correct, every twin pair discriminated — the thinking tokens spent fall by about half, and by roughly three quarters measured from the oldest perfect model in the ladder. That is a real efficiency gain and it would be unfair to read this table as “thinking is useless”. What the table shows is narrower and more exact: after the ceiling, reasoning has no marginal utility on this function. Half the thinking still buys the same verdicts as twice the thinking, and both still lose on latency to a model that does no thinking at all. What reasoning is worth on other tasks is not measured here.

Once the contract is narrow enough, model progress can continue while task utility has already saturated. The interesting number a benchmark can produce is not which model is best — it is the point after which additional capability stops having any value for the operation being bought.
Judgement SLA

Does it read the evidence correctly?

Unsupported closure, paired discrimination, over-abstention. A property of the model, stable across repeated draws for a local model, noisy by about one case for a hosted one.

Serving SLA

Will it answer before the deadline?

p50, p95, max, and tail ratio. A property of the service, not the model. One hosted model answered with a 1.19s median and one 256s stall — 92% of that run's total wall in a single request. A mean would have hidden it; a total would have made the model look 200× slower than it is.

11 · Anthropic on the lift

The elevator passed the useful floor and kept going

Ten models across three classes and several generations. Three older IDs the operator asked for are retired and are recorded UNAVAILABLE rather than replaced by a nearby alias. Every row also stores the model the response itself reports, because the top-tier model can fall back to another one and crediting it for an answer something else gave would be the exact fabrication this bench exists to prevent.

Eight of the ten score a perfect zero, including the cheapest and fastest model in the lineup. The two most capable models are the only ones that fail — and they fail by declining to answer.

This does not repeat the previous vendor’s curve, it continues it. There, the cheap tier reached the ceiling and everything above it scaled capability at flat utility. Here the cheap tier reaches the same ceiling, capability scales flat for eight more models — and then, in the newest generation, utility bends downward, not because judgement degraded but because a policy started withholding it.

The refusals are not spread evenly. They fall on the adversarial cases — the ones that test whether a verifier catches a false claim — while every defect-removed positive twin is answered normally. A verifier that answers each easy case and goes quiet on the traps has a coverage hole shaped exactly like the threat it was hired to detect.

Re-running the refused cases three times each shows the behaviour is mostly content-deterministic: four cases refuse every time, one answers every time, one flips. So the refusal count is draw-dependent even though the phenomenon reproduces. The cause is not established — the obvious hypothesis, that security vocabulary trips a classifier, is falsified by a positive twin containing the same words and answering fine.

And no fallback occurred: the provenance census returned the requested model on all twenty cases, so the score belongs to the model that was asked.

One caveat against reading these two as excellent judges. Their accuracy on answered cases is 100%, but that figure is survivorship: the cases they decline are the hard ones. When a different draw forced an answer on a case it had previously refused, the answer was wrong. What remains true, and matters most, is that their unsupported closure is zero — every failure is a missing verdict, never a closure that the evidence did not support. For a tier-0 verifier that is the safe direction to fail in; it is simply not a usable one.

12 · The effort knob

More reasoning is billing, until it is damage

Four models chosen as a phenotype ladder — one with no refusals and no effort control, one at ceiling, one mildly refusing, one strongly refusing — crossed with every effort level the surface accepts.

A model already at ceiling gains nothing from effort at any level. It spends six times the tokens to return the same verdicts, and at the top setting it spends a hundred and twenty-five times as many to return fewer.

Three different mechanisms, one outcome. A model can fail to deliver a verdict by abstaining on epistemic grounds, by declining on policy grounds, or by reasoning past its output budget without ever concluding. Those are unrelated causes with identical operational consequences, so decision coverage counts all three and keeps the causes separable. Loss alone cannot tell them apart, and a verifier that never answers is not safer than one that answers correctly — it is just absent.

Higher capability tier does not imply higher decision coverage. Across this ladder it implies the opposite.

The conditional-accuracy column is a trap worth reading carefully. For the strongest refuser it reads exactly 1.000 in every configuration where it declines six cases, and falls below that in every configuration where it declines fewer. Its perfect record on answered cases was its refusals removing its own hard cases. Accuracy conditional on answering flatters any model that gets to choose which questions count.

Two properties survive all of it. Refusal is effort-invariant for the mild phenotype — the same two cases at every level, nine times the tokens between the cheapest and dearest setting — so it is a policy property, not a reasoning-budget artifact. And through every model and every effort level in this arm, unsupported closure never left zero: these models fail by silence, never by closing something the evidence did not support. For a tier-0 verifier that is the right direction to fail in. It is still a failure.

13 · A third vendor

Three generations of progress. Same verdicts. More waiting.

By this point correctness has nowhere left to go: several models already hold every case and every twin pair. A newer flagship can only match that ceiling more cheaply, match it more expensively, or find a new way to miss it.

Across three consecutive generations the verdicts are identical and everything else gets worse: four and a half times the reasoning tokens, nearly three times the median latency, and eleven times the tail.

This vendor did not find a new way to fail, which is itself worth recording: decision coverage is 1.00 everywhere, with no refusals and no truncations. It always answers. What it does have is the worst tail in the study — a median around ten seconds and a slowest case twenty-six times that. Under a mean, that model looks fine.

The most interesting entry is a retired checkpoint published in two variants that differ only in whether it reasons. The non-reasoning half is the fastest model in this arm and discriminates only one twin pair in four; the reasoning half takes ten times as long and discriminates all four. That is the cleanest evidence in the whole study that reasoning buys discrimination — and it sits on a deprecated model, not a flagship.

Reasoning can help. This is the checkpoint where it actually did.

That is the honest shape of the reasoning result across the whole study, and it is neither slogan. Extra deliberation is not a quality dial that turns the same way for everybody; it is an intervention whose effect is specific to the task, the family and the individual checkpoint. In this suite the same switch produced repair on one model, partial and expensive repair on another, no benefit at all on a third, and outright non-termination on a fourth. A benchmark that could only report reasoning as good, or only ever as waste, would have been unable to see any of that.

14 · The last ladder

We kept buying more thinking until it started making the answers worse

The cleanest control available: one vendor, one generation, three capability tiers, and an explicit reasoning dial, run over the same frozen evidence and scored by the same scorer as everything above. The published top setting of that dial does not exist — every tier rejects it — so it is recorded as unavailable rather than quietly folded into the setting below it.

In all three tiers the best cell is the one with reasoning switched off. Two of them score a clean sweep at zero reasoning tokens, and the top tier's worst cell is its highest effort setting.

Read the first column on its own and the product ladder works exactly as advertised: the cheapest tier misses a case, the two above it do not. Read across the rows and the dial that is supposed to buy judgement is buying the opposite. That result was too convenient to admit on one draw each, so the two endpoints were re-run three times.

The repeats say something the single draws could not. With reasoning off the verdicts did not move across these draws — the cheap tier missed the same case all three times, the top tier was clean all three times. Switch reasoning on and the failures start wandering: three draws at the highest effort failed on three different sets of cases. Three draws cannot establish that a hosted backend is deterministic, and this does not claim it. What it does show is that the configuration spending nothing was the steadier of the two, while the configuration spending the most produced answers that changed between identical requests without ever getting better.

What the arm never does is close a clause it should not have. Across every cell, unsupported closure is zero, decision coverage is 1.00, and there are no refusals and no truncations. Every error in the whole factorial is the same one: declining to commit on the two hardest adversarial twins, where the discriminating fact is an ordering between a timestamp and a revision. That is the safe direction to be wrong in, and it is priced accordingly.

Two of the three cases that ever move share a shape worth naming. Neither asks whether evidence is present; both ask for a relation between two facts — whether a test predates the implementation it claims to cover, whether a receipt names the revision under review. That is a harder operation than classifying a single artifact, and it is where every tier of this ladder reaches for the non-committal answer. Whether ordered and bound comparison is genuinely a weaker class, or these two cases are merely awkwardly worded, is not something twenty cases can settle. It is the next experiment, not a result of this one.

The cost axis closes the study. This entire factorial — three tiers, five effort settings, several hundred judgements — spent fewer reasoning tokens in total than a single competing flagship spent on twenty. And the cheapest cells in the grid, the ones that spent none at all, hold two of its three perfect scores.

On this frozen contract, additional reasoning effort bought no systematic improvement at any tier. The reasoning-off configuration was the best observed cell in all three, and repeated runs of the top tier scored lower, and less stably, at the highest effort than at none.

The cheapest successful computation was no reasoning at all. A roleplay fine-tune had already answered, in 672 milliseconds, on a nine-year-old consumer graphics card.

15 · Pedigree

What a model is called predicts the odds, not the answer

Every model tagged by what its name advertises — flagship, cheap tier, coder, vision-language, roleplay merge — because that is the information available when choosing a verifier from a model card, before running anything.

The last vendor to arrive had to be tagged general, three times over, and that is not a gap in the tagging. Its three tiers are named after a moon, a planet and a star; nothing in those identifiers says which is the cheap one. The newest ladder in the study carries no readable pedigree at all — there is no model-card claim left to be right or wrong about.

All three of them miss the gate here, which requires both zero unsupported closures and all four twin pairs separated. They miss it on the pairs, at their shipped default settings — and two of the three clear it comfortably once the reasoning dial is turned off. Measured the way every other model on this page was measured, at whatever the vendor ships as the default, the newest tiers from the largest vendor do not qualify.

The label is not worthless. Models whose names advertise flagship status clear the gate about two and a half times as often as roleplay merges. Anyone claiming pedigree carries no information is arguing with this table.

But it predicts the odds of clearing the bar, not which model to pick once several have. The fastest perfect score in the entire study belongs to an uncensored roleplay merge, and the first flagship appears seventh.

The last two rows are the two vendors with nothing in the eight above, each shown in the configuration where it actually clears the gate. Neither earns its place on speed: the first does not qualify at all at the setting it ships, and the second is five times slower than the row above it. They are there to be seen, not ranked.

Both readings are true at once and the second is the one that changes a decision. If candidates had been shortlisted the sensible way — by reading model cards — the roleplay swamp would have been discarded first, and with it the fastest sufficient verifier found here. Model-card plausibility is not an admission criterion. The benchmark does not care what a model was meant to be. It only reports what the model did with the evidence it was given.

Qualification attaches to the artifact, not to the dignity of its name.
16 · Limits

What this does not establish