A closure-contract verifier answers one question: does this evidence support closing this clause? Getting it wrong in the permissive direction promotes unsupported state. We ran the identical frozen suite from a 230M model on a 2017 GPU up to a 2.81T hosted frontier model, to find where sufficiency actually begins.
Three were found by the operator noticing something that looked wrong; two were self-inflicted mid-run. Several produced no error at all and broke nothing visible, which is what made them expensive.
Both local models reported exactly 58s. The client addressed the
provider as localhost, which resolves to ::1 first; the
refused connection cost ~2.05s per request. Constant across a 24×
change in prompt size, so not prefill. 79% of the published latency was socket
floor. Verdicts were unaffected — which is what made it invisible.
The parser took the last status token anywhere in the reply. A reasoning model whose thinking never terminated, musing “I think it is SATISFIED”, parsed as SATISFIED — a closure manufactured from a model that never issued one. Note the direction: it could only ever fabricate closure, never refuse one, matching the failure it was built to hunt.
Two identical runs disagreed on 12 of 20 cases for one model and 0–1 for most others. The local control reproduced exactly (6, 6, 6 — zero flips), so the variance is the provider, not the harness. Any cloud difference of one to two cases is inside the noise.
A follow-on job waited for the results file to stop growing instead of asking whether the process was alive. A model taking 38s per verdict stalled the file past the threshold, so a second run launched onto the same GPU — each unloading the other’s resident model. 17 records were destroyed. Seven came back as HTTP 400 and read as “model incompatible”; retested individually afterwards, all of them loaded and answered fine. Same shape as R2: a proxy standing in for the thing itself.
Three calls returned HTTP 429. The scorer treated a row with no
answer as UNPARSEABLE and charged 20 loss points, so two models were published at
loss 20 whose real loss was 0 — and it manufactured a tidy
false narrative that the newer generation was regressing. The scorer now excludes
rows the model never answered, and reports how many cases the metrics rest on.
Unevaluated is not failed.
Three of the four could only ever bias results in a direction that flattered the run — a latency floor that hid the cheap models, a parser that manufactured closure, and a corruption that looked like model incapacity. An instrument whose failures all point one way cannot be trusted to measure that direction, which is why they are reported here rather than fixed quietly.
Each point is one local model. Vertical position is unsupported closure — how many times it closed a clause the evidence did not support, the one error the architecture exists to prevent.
The same cloud models, run twice, with provider-side reasoning as the only variable. Lines falling to the left are models rescued by reasoning.
One instance was withdrawn. An apparent 207 → 6 repair
turned out to be a bad draw: that model scores 7, 3, 207 across repeated identical runs,
and its thinking-on score of 6 is simply its normal score. The effect survives where the
baseline is stable — one model scored 15 three times running before dropping to 0
with reasoning, and the local olmo-3-7b instruct/think pair is deterministic.
Differences of one to two cases sit inside the measured noise floor, so the frontier band is not rank-ordered by this data. What is established is the absence of benefit, not the presence of harm: nothing in the pool beats a 20B, and a local 12B on a 2017 GPU matches the best hosted result exactly. The largest model tested carries 2.81T parameters by its published spec — roughly 140× that 20B and 234× the local 12B — and scores the same.
Gate: zero unsupported closure and 4/4 paired discrimination. The paired test uses adversarial cases and their defect-removed twins, so a model that pattern-matches suspicion fails it, and so does one that is merely timid.
The table above is a list of passes, so on its own it hides the thing this contract exists to catch. These are the runs that closed a clause the evidence did not support — the error that promotes unsupported state, and the reason the loss function prices it a hundred times higher than being wrong the other way.
Counted over every complete run in the study, on the cases where closing is never the admissible answer. Runs, not distinct models: the pool holds sibling checkpoints, quantisation variants and repeat configurations of the same line, so these are counts of measured artifacts. If unsupported closure were scattered model-specific error, the rates would be roughly level across cases. They are not.
Three unrelated models from three different families make exactly this jump, which is what separates a shared attractor from scattered noise. It also settles a tempting shortcut: one of them separates all four twin pairs and still closes an unsupported clause here. Telling a defect from its repaired twin and refusing to over-claim are different abilities, and a gate that tests only the first would have passed it.
Several cases admit more than one correct answer, so a verdict can change between identical requests and cost nothing at all. Counted that way, two models below look similarly noisy. Split by whether the change leaves the admissible set, they are not alike in the slightest.
The distinction is the difference between a judge who words the same refusal differently and a judge who refuses on Monday and permits on Tuesday, given the identical file. Only the second is a hazard, and the raw flip count cannot tell them apart.
What counts as harmless here is a property of this contract, not of the words. Two of the five statuses are treated as interchangeable because under these rules both refuse to close, so either one leaves the state exactly where it was. A system that gave them different consequences — one asking for more evidence, the other ending the attempt — would have to redraw the line, and some of the drift counted as free above would stop being free. The unit to measure is not whether the answer changed, but whether what the answer authorises changed.
Local models reproduce exactly here — the control returned 6, 6, 6 with zero flipped cases — so a difference between two builds of one model is a difference, not a draw.
It is not a “more bits is better” law, which is what makes it awkward: two families got worse with a larger quantisation. The effect is real, model-specific and non-monotonic, so it cannot be predicted — only measured.
A separate arm, run through a hosted CLI verifier harness. It is not on the same axis as the other two: the harness adds its own read-only-verifier framing and returns a structured object instead of a bare status word. Provenance was confirmed on every call — the route reports which model actually answered, and reports nothing when none did.
The heavier tier scored worse than the lighter one, and every one of its errors ran the same way: it refused to close cases the oracle calls plainly satisfied — including the control case that exists to catch exactly that. Safe, but a verifier that can never say yes escalates everything.
One model in this arm was also scored on the cloud arm, as a check on the harness itself. It moved by a single case — which is precisely the measured noise floor, so the harness effect is not established and is indistinguishable from a draw at one run.
One checkpoint, four ways of asking for the same verdict, three repeated draws each. Same provider, same credential, same frozen evidence, same scorer. The only thing that varies is the shape the answer must arrive in.
This kills two comfortable explanations at once. It is not that the nested shape overloads the model — the minimal one-field version produced the single worst draw in the experiment. And it is not that the model needs room to think out loud before committing — the arm with no room at all, one word and nothing else, was the best of the four.
What the constrained arms lost was reproducibility, not one identifiable ability. But the raw count of unstable cases hides the distinction that decides whether instability costs anything. Several cases admit more than one correct answer, so a verdict can move between draws and stay entirely within the admissible set. That kind of drift is free. The kind that matters moves across the boundary — and the two behave nothing alike here.
So the one-word arm is not bit-stable, and this does not claim it is: it answered two cases differently across draws. It simply never used that freedom to leave the region where any answer is acceptable. The structured arms did, repeatedly. And they did it in much the same places: six of the eight cases that drift do so under both of them. All eight were already weak surfaces by an independent measure — every one of them had drawn false closures elsewhere in that census, at rates from to , and not one was a case the census found clean. Runs, again, not distinct models: that pool holds sibling checkpoints and quantisation variants of the same lines. So within this frozen suite the imposed format loosened surfaces that were already measurably weak rather than opening a new kind of hole. Outside these twenty cases it could do either, and nothing here speaks to that.
That is worse than malformed output, which at least announces itself. Here the contract the machine checks was satisfied completely and the contract that matters was not. Note what is and is not established: the split in this experiment falls exactly on whether a format was imposed. What happens behind that switch — constrained decoding, a separate provider path, a prompt transformation, something else — was not observed and is not claimed. Three of the four models tried elsewhere showed no such effect at all, so this is one checkpoint's behaviour, not a law.
The engineering reading survives that caution intact. The imposed format was trying to enforce, during stochastic generation, a property the runtime can check deterministically once generation is over. A reply either is one of five permitted words or it is not, and answering that question is one comparison. Doing it afterwards does not make the model more likely to comply — it makes non-compliance cheap, visible, and unable to reach the state it would have changed.
Ten fixed model IDs from one provider — no -latest
aliases, because qualification binds to a configuration and an alias can be repointed at
another checkpoint. Every record carries the provider's own reported version, and all ten
matched the ID requested.
The cheapest tier posts a perfect score with zero thinking tokens, while the most expensive tier spends tens of thousands of them, takes roughly 28× as long, and scores worse. Stated precisely: this is not a leaderboard win. On raw median latency the cheapest tier is actually beaten by 16 ms — by a later model with a worse score. The claim that survives a literal reading is about the ceiling: it was reached two generations ago, and nothing since has exceeded it. The only genuine regression is inside the cheap line itself, where a newer Lite is worse than the one it replaced.
The reasoning budget is genuinely getting cheaper. Across the three newest models that all score identically — every case correct, every twin pair discriminated — the thinking tokens spent fall by about half, and by roughly three quarters measured from the oldest perfect model in the ladder. That is a real efficiency gain and it would be unfair to read this table as “thinking is useless”. What the table shows is narrower and more exact: after the ceiling, reasoning has no marginal utility on this function. Half the thinking still buys the same verdicts as twice the thinking, and both still lose on latency to a model that does no thinking at all. What reasoning is worth on other tasks is not measured here.
Unsupported closure, paired discrimination, over-abstention. A property of the model, stable across repeated draws for a local model, noisy by about one case for a hosted one.
p50, p95, max, and tail ratio. A property of the service, not the model. One hosted model answered with a 1.19s median and one 256s stall — 92% of that run's total wall in a single request. A mean would have hidden it; a total would have made the model look 200× slower than it is.
Ten models across three classes and several generations. Three older IDs the operator asked for are retired and are recorded UNAVAILABLE rather than replaced by a nearby alias. Every row also stores the model the response itself reports, because the top-tier model can fall back to another one and crediting it for an answer something else gave would be the exact fabrication this bench exists to prevent.
This does not repeat the previous vendor’s curve, it continues it. There, the cheap tier reached the ceiling and everything above it scaled capability at flat utility. Here the cheap tier reaches the same ceiling, capability scales flat for eight more models — and then, in the newest generation, utility bends downward, not because judgement degraded but because a policy started withholding it.
The refusals are not spread evenly. They fall on the adversarial cases — the ones that test whether a verifier catches a false claim — while every defect-removed positive twin is answered normally. A verifier that answers each easy case and goes quiet on the traps has a coverage hole shaped exactly like the threat it was hired to detect.
Re-running the refused cases three times each shows the behaviour is mostly content-deterministic: four cases refuse every time, one answers every time, one flips. So the refusal count is draw-dependent even though the phenomenon reproduces. The cause is not established — the obvious hypothesis, that security vocabulary trips a classifier, is falsified by a positive twin containing the same words and answering fine.
One caveat against reading these two as excellent judges. Their accuracy on answered cases is 100%, but that figure is survivorship: the cases they decline are the hard ones. When a different draw forced an answer on a case it had previously refused, the answer was wrong. What remains true, and matters most, is that their unsupported closure is zero — every failure is a missing verdict, never a closure that the evidence did not support. For a tier-0 verifier that is the safe direction to fail in; it is simply not a usable one.
Four models chosen as a phenotype ladder — one with no refusals and no effort control, one at ceiling, one mildly refusing, one strongly refusing — crossed with every effort level the surface accepts.
Three different mechanisms, one outcome. A model can fail to deliver a verdict by abstaining on epistemic grounds, by declining on policy grounds, or by reasoning past its output budget without ever concluding. Those are unrelated causes with identical operational consequences, so decision coverage counts all three and keeps the causes separable. Loss alone cannot tell them apart, and a verifier that never answers is not safer than one that answers correctly — it is just absent.
The conditional-accuracy column is a trap worth reading carefully. For the strongest refuser it reads exactly 1.000 in every configuration where it declines six cases, and falls below that in every configuration where it declines fewer. Its perfect record on answered cases was its refusals removing its own hard cases. Accuracy conditional on answering flatters any model that gets to choose which questions count.
Two properties survive all of it. Refusal is effort-invariant for the mild phenotype — the same two cases at every level, nine times the tokens between the cheapest and dearest setting — so it is a policy property, not a reasoning-budget artifact. And through every model and every effort level in this arm, unsupported closure never left zero: these models fail by silence, never by closing something the evidence did not support. For a tier-0 verifier that is the right direction to fail in. It is still a failure.
By this point correctness has nowhere left to go: several models already hold every case and every twin pair. A newer flagship can only match that ceiling more cheaply, match it more expensively, or find a new way to miss it.
This vendor did not find a new way to fail, which is itself worth recording: decision coverage is 1.00 everywhere, with no refusals and no truncations. It always answers. What it does have is the worst tail in the study — a median around ten seconds and a slowest case twenty-six times that. Under a mean, that model looks fine.
The most interesting entry is a retired checkpoint published in two variants that differ only in whether it reasons. The non-reasoning half is the fastest model in this arm and discriminates only one twin pair in four; the reasoning half takes ten times as long and discriminates all four. That is the cleanest evidence in the whole study that reasoning buys discrimination — and it sits on a deprecated model, not a flagship.
That is the honest shape of the reasoning result across the whole study, and it is neither slogan. Extra deliberation is not a quality dial that turns the same way for everybody; it is an intervention whose effect is specific to the task, the family and the individual checkpoint. In this suite the same switch produced repair on one model, partial and expensive repair on another, no benefit at all on a third, and outright non-termination on a fourth. A benchmark that could only report reasoning as good, or only ever as waste, would have been unable to see any of that.
The cleanest control available: one vendor, one generation, three capability tiers, and an explicit reasoning dial, run over the same frozen evidence and scored by the same scorer as everything above. The published top setting of that dial does not exist — every tier rejects it — so it is recorded as unavailable rather than quietly folded into the setting below it.
Read the first column on its own and the product ladder works exactly as advertised: the cheapest tier misses a case, the two above it do not. Read across the rows and the dial that is supposed to buy judgement is buying the opposite. That result was too convenient to admit on one draw each, so the two endpoints were re-run three times.
The repeats say something the single draws could not. With reasoning off the verdicts did not move across these draws — the cheap tier missed the same case all three times, the top tier was clean all three times. Switch reasoning on and the failures start wandering: three draws at the highest effort failed on three different sets of cases. Three draws cannot establish that a hosted backend is deterministic, and this does not claim it. What it does show is that the configuration spending nothing was the steadier of the two, while the configuration spending the most produced answers that changed between identical requests without ever getting better.
What the arm never does is close a clause it should not have. Across every cell, unsupported closure is zero, decision coverage is 1.00, and there are no refusals and no truncations. Every error in the whole factorial is the same one: declining to commit on the two hardest adversarial twins, where the discriminating fact is an ordering between a timestamp and a revision. That is the safe direction to be wrong in, and it is priced accordingly.
Two of the three cases that ever move share a shape worth naming. Neither asks whether evidence is present; both ask for a relation between two facts — whether a test predates the implementation it claims to cover, whether a receipt names the revision under review. That is a harder operation than classifying a single artifact, and it is where every tier of this ladder reaches for the non-committal answer. Whether ordered and bound comparison is genuinely a weaker class, or these two cases are merely awkwardly worded, is not something twenty cases can settle. It is the next experiment, not a result of this one.
The cost axis closes the study. This entire factorial — three tiers, five effort settings, several hundred judgements — spent fewer reasoning tokens in total than a single competing flagship spent on twenty. And the cheapest cells in the grid, the ones that spent none at all, hold two of its three perfect scores.
The cheapest successful computation was no reasoning at all. A roleplay fine-tune had already answered, in 672 milliseconds, on a nine-year-old consumer graphics card.
Every model tagged by what its name advertises — flagship, cheap tier, coder, vision-language, roleplay merge — because that is the information available when choosing a verifier from a model card, before running anything.
The last vendor to arrive had to be tagged general, three times over, and that is not a gap in the tagging. Its three tiers are named after a moon, a planet and a star; nothing in those identifiers says which is the cheap one. The newest ladder in the study carries no readable pedigree at all — there is no model-card claim left to be right or wrong about.
All three of them miss the gate here, which requires both zero unsupported closures and all four twin pairs separated. They miss it on the pairs, at their shipped default settings — and two of the three clear it comfortably once the reasoning dial is turned off. Measured the way every other model on this page was measured, at whatever the vendor ships as the default, the newest tiers from the largest vendor do not qualify.
The label is not worthless. Models whose names advertise flagship status clear the gate about two and a half times as often as roleplay merges. Anyone claiming pedigree carries no information is arguing with this table.
The last two rows are the two vendors with nothing in the eight above, each shown in the configuration where it actually clears the gate. Neither earns its place on speed: the first does not qualify at all at the setting it ships, and the second is five times slower than the row above it. They are there to be seen, not ranked.
Both readings are true at once and the second is the one that changes a decision. If candidates had been shortlisted the sensible way — by reading model cards — the roleplay swamp would have been discarded first, and with it the fastest sufficient verifier found here. Model-card plausibility is not an admission criterion. The benchmark does not care what a model was meant to be. It only reports what the model did with the evidence it was given.
Help move GitDuck Bench from GTX 1080 Ti to B300.
The ducks are ready. The tensor cores are not. ๐ฆ