What happens to a set of benchmark results when the same submissions are scored a second time, shuffled and blind, under a matched effort protocol? Nine of ten moved. All but one moved up, the largest by 27 points, and the error was not random. It tracked how bad each submission looked.
Nothing about the results. A later run happened to log its grader activity, and two cells of the same size had been graded in 9 tool calls and 58.
Both graders returned a confident number. Neither number was marked as less reliable than the other. That disparity said nothing about which score was right, but it did say the amount of reading behind a score was uncontrolled, and it had been uncontrolled for the entire preceding series.
Every claim in the lab rested on those numbers. So the whole set went back through.
The lab has one procedural rule that matters: the hypothesis goes in the notebook before the run. This run broke it. The predictions were stated out loud while setting the re-grade up and never written down first, which makes them weaker evidence than they look, and I am recording that rather than presenting them as if they had been pre-registered.
What was claimed, in conversation, beforehand: that the repeat run would confirm a small noise floor, that the instrument's saturation at the top of the range would survive, and that one flat result at the middle tier needed checking.
The expectation underneath all three was that the re-grade would be housekeeping. Small movements, no changed conclusions, cheap insurance.
All ten submissions copied into a single shuffled, interleaved pool with every prior score file stripped out. The originals had been graded blind too, but in batches labelled by condition, which tells a grader what it is looking at even when it does not know the score.
Ten cold graders, one per cell, each given a mandatory protocol: enumerate the files, read all of them, score item by item, and report the file-read count. The count is the control. A grader that reports reading four files out of thirty has disqualified its own number.
| Cell | Original | Re-grade | Δ |
|---|---|---|---|
| Strong tier, full library | 96 | 100 | +4 |
| Strong tier, full library (repeat run) | 100 | 100 | 0 |
| Strong tier, thin library | 93 | 100 | +7 |
| Strong tier, colloquial library | 100 | 100 | 0 |
| Strong tier, no library | 62 | 89 | +27 |
| Mid tier, full library | 97 | 97 | 0 |
| Mid tier, thin library | 97 | 97 | 0 |
| Mid tier, no library | 89 | 92.5 | +3.5 |
| Low tier, canonical library | 78 | 82 | +4 |
| Low tier, colloquial library | 84 | 81 | −3 |
Nine of ten scored the same or higher. The one that fell, fell by 3.
Random grader variance would scatter in both directions with roughly equal magnitude. This did not. Nine of ten corrections went one way, and the size of each correction tracked how bad the submission looked on first inspection.
The 27-point miss landed on the lowest-scoring cell in the series. The zero-movement cells were the ones that already looked polished. A submission that presented badly was read less carefully and scored more harshly, and the two compounded.
That is presentation bias contaminating a substance score. It matters here because the whole point of the instrument was to separate looks right from is right, and the grader was doing the opposite at exactly the cells where the distinction mattered most.
The original scoring was already blind. Blinding was not enough, because the graders were free to spend wildly different amounts of attention per submission. Blind grading with unmatched effort is not a controlled comparison. The fix that worked was cheap: a mandatory read protocol, and a reported read count that makes an under-read score visible.
A phantom noise floor. Two independent cold runs of the identical input, same model and same context, had scored 96 and 100. That 4-point spread had been treated all afternoon as the instrument's noise floor, and used to demote findings smaller than 4 points as unreliable.
Under matched grading the two runs agree exactly, dimension for dimension, both at 100. The noise floor is zero. Generation on this task is far more reproducible than the series implied, and every finding demoted for falling under 4 points had been demoted against a number that did not exist.
A phrasing effect. A 6-point gap between two differently-worded versions of the same library closed to 1 point and changed sign. It was variance in two specific cells. The decision that gap had driven, which was to stop trusting a set of vocabulary rules, still stands. The reason given for it was wrong.
That distinction is worth keeping separate. Right call, wrong reason, is a different state from right call, right reason, and it stays fragile until the reason is fixed.
A context effect, by two thirds. The headline result from the ablation study, that removing a pattern library cost 34 points, became 11. The direction survived. The magnitude did not, and a pre-registered prediction that had been recorded as confirmed turned out to be wrong by more than twenty points. That one has its own write-up.
One time-boxed block and ten cold graders. Against that: a day of conclusions that had to be re-derived, two findings withdrawn or rewritten, and every number in the repository stamped with a warning pointing at the corrected series.
The re-grade itself was the cheapest thing in the sequence. The expensive part was everything built on top of the numbers before anyone checked them.
The broader habit this reinforced: an evaluation harness is itself an artifact that can be wrong, and it is the one artifact nobody tests, because it is what testing is done with. Six of ten predictions in this lab have come back wrong so far. Until this run, none of the misses had been attributed to the instrument.
The frozen spec, rubric, all ten submissions, the original and
re-graded score files, and the re-grade protocol live in the lab
repository. Superseded pages carry warnings pointing at the
corrected series rather than having been rewritten in place.
Related: a pattern library buys conventions, not judgment