For months, developers have claimed that frontier models get quietly worse in the weeks after they ship. The counter-claim is that people are pattern-matching on noise. Both sides have been arguing without an instrument, because nobody kept a clean record of how the model behaved on day zero.
A repository called livenerf, started the day Claude Opus 5.5 shipped, is the first serious attempt to fix that. Before it collected a single day of series data, it tested its own instrument against three deliberate, known degradations. On accuracy, it caught none of them.
TL;DR
- livenerf pre-registered a 30-day benchmark to detect whether Claude Opus 5.5 degrades after launch, then published its own blind spot: swapping in a previous-generation model was not distinguishable at 99% confidence.
- Three deliberate degradations were measured. On accuracy, none cleared the bar: z-scores of 1.08, 1.84 and 0.60 against a threshold of 2.58. On output token counts, the same samples gave 6.6, 9.9 and 2.0.
- Running the identical model against itself produced a +6.4 point accuracy swing, larger than two of the three real degradations. The noise was louder than the signal.
- The practical takeaway for anyone running AI in production: accuracy evals are a statistically expensive way to detect supplier-side change, and your token counts are already in the response body for free.
What the repository actually does
livenerf runs a frozen panel of benchmark questions once a day for 30 days and compares each day against a launch-week baseline. It is built on Inspect, the UK AI Security Institute’s eval framework, and its statistics follow Evan Miller’s 2024 paper Adding Error Bars to Evals. The design is deliberately boring: frozen prompts, a pinned CLI, exact-match grading, raw logs kept forever.
One design note is worth lifting out, because it contradicts advice we have given on this site. The repository refuses to use an LLM as a judge, and says why: “no LLM judge, ever, since the judge would drift too.” If you are measuring drift, you cannot grade with something that drifts. Our own 2026 guide to LLM observability recommended graduating to LLM-as-judge evaluations for semantic quality. That advice holds for measuring your own application. It does not hold for measuring your supplier.
The pre-registration is real, and the margins are thin
The repository’s central credibility claim is that its protocol was committed to public git before the data it governs existed. That is checkable, so we checked it against the GitHub API rather than taking it on trust.
It holds. The repository was created at 21:59 UTC on 22 September 2026, the day of the launch it measures. The protocol v2 commit landed at 02:08:17 UTC on 24 September; the confirmation samples it governs ran from 02:49 that morning, a clear 41-minute margin. Two other orderings are tighter than they look: the design-lock commit is timestamped 04:41:50 UTC against a validation run logged as starting at 04:41, and the day-one commit is 22:10:11 UTC against a series start logged as 22:10. Those two are consistent with the claim but cannot be independently resolved from public data, because the logs record minutes and the commits record seconds. The deviations log is the more interesting artefact anyway: seven amendments, six of them logged before the baseline started, each dated and reasoned. Pre-registration did not make the protocol stable. It made the instability auditable, which is the part worth copying.
Three degradations, and what caught them
Before trusting any null result, the author ran a positive control: degrade the model in known ways and check the rig notices. Three perturbations were tested on the same 78-item panel.
| Deliberate change | Accuracy Δ | Accuracy z | Output tokens | Token z |
|---|---|---|---|---|
| Effort medium instead of high | −4.2 ± 3.9 | 1.08 | −26% | 6.6 |
| Effort low instead of high | −8.3 ± 4.5 | 1.84 | −62% | 9.9 |
| Previous-generation model swapped in | −3.8 ± 6.3 | 0.60 | −23% | 2.0 |
A two-sided 99% test needs a z-score of 2.58. The accuracy column does not reach it once. The token column clears it twice and comes close on the third. On identical samples, the token signal is between 3.3 and 6.1 times more statistically significant than the accuracy signal.
This is the finding, and it is not a criticism of the rig. It is a property of what is being measured. A model that thinks less produces fewer tokens immediately and obviously; whether it also gets answers wrong is a noisy, second-order consequence.
The control that should end the argument
The most quotable number in the repository is not any of the degradations. It is the A/A check: the same model, same effort, same panel, split into two halves and compared against itself. That comparison returned +6.4 points, standard error 3.6.
Nothing changed, and accuracy moved 6.4 points, in the flattering direction, and by more than the measured effect of either the model swap (3.8) or the drop to medium effort (4.2). Every “the model feels worse this week” thread on the internet is arguing about a quantity smaller than this rig’s own null result.
Why accuracy is such a weak instrument
The calibration record explains the underlying problem. The author screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition maths and AIME with four samples each. Of the 2,332 fully screened, 2,119 were always right and 133 were always wrong. 96.57% of a standard benchmark pool carried no information about drift whatsoever.
That left 80 eligible questions, and a final panel of 78, or 3.3% of the pool. A report-only audit of the panel found 8 answer keys that look wrong and 30 questions that are ambiguous, which is 48.7% of the items the whole measurement rests on. None of that is sloppiness. It is what the informative tail of a benchmark looks like when the model is already at 93% on the first try.
The cost of fixing this with more samples is published too, and it is brutal. The series spends 3.6 points of the plan’s weekly usage meter to reach a minimum detectable effect of 7.5 accuracy points per 10-day window. Buying an MDE of 4 points costs 11.3 weekly-meter points: 3.1 times the actual spend, and over the 10-point cap the design set for itself. The token signal cost nothing extra: those counts arrive with samples that were already being taken.
What this means if you ship on someone else’s model
We have written before about the model inside your tooling arriving under a contract you are not a party to, and about how, when you self-host, the serving stack becomes part of the model. This is the third case, and the least tractable: a hosted model where the serving stack is not yours to pin, the contract stays intact, and the thing that changes is the output.
Note that substitution is not hypothetical here. livenerf documents that the serving path sometimes answers with a different model entirely, and rejects those samples as classifier events. The question was never whether a served model can change. It is whether you would know.
Four things follow for engineering leaders:
- Log output tokens per request, and alert on the distribution, not the mean. Every provider returns a usage object. This is the cheapest drift canary available and the evidence says it is the most sensitive one.
- Do not commission an accuracy eval to police your supplier. To resolve a change the size of a model swap you need roughly an order of magnitude more sampling than most teams will fund, and your result will still be swamped by run-to-run noise.
- Pin what is actually pinnable, and record it. The harness is yours: CLI version, system prompt, tool set, effort level. livenerf’s rule is the right one, because a changed harness looks exactly like a changed model. Most teams upgrading their agent tooling weekly have destroyed their own baseline without noticing.
- Take the baseline on day one. The reason this argument has run for months without resolution is that nobody had a before. Capturing token distributions and pass rates in the first week of adopting a model costs almost nothing and is unrecoverable later.
One caveat the repository is scrupulous about and we will repeat: none of this is a drift result. The series is 6 of 30 days in, the baseline is still being collected, and the first decision cannot arrive before late October. Anyone citing livenerf today as proof that a model was or was not nerfed is citing a rig, not a finding.
REPTILEHAUS builds token-level observability, evaluation harnesses and the DevOps work to pin a harness properly. If you are running a model you do not control and have no baseline for it, get in touch.
Figures verified against ninjahawk/livenerf on 30 September 2026: the README, PREREGISTRATION.md and the docs/ directory, with commit timestamps from the GitHub API. The series is live and the repository is updated daily; re-check the source before quoting these numbers.
📷 Photo by Ag PIC on Unsplash


