When an Unchanged Model Looks Improved
A per-problem benchmark transition is evidence about a model only after the evaluation procedure's own no-change behavior is understood; otherwise inference nondeterminism, finite sampling, grading, and thresholds can be mistaken for learning or forgetting.
A language model was evaluated on the same math problems twice. The second time, it seemed to have learned six problems and forgotten nine.
Nothing had been done to teach it.
The model was the same Qwen3-8B, with the same frozen weights. The apparent changes came from a ledger that gave each problem one greedy answer, called the answer solved or unsolved, and compared the two runs. Six entries moved one way. Nine moved the other.
That is the useful puzzle in Phantom Gains: Auditing Self-Improvement Against a Measured Null, an arXiv v1 preprint submitted on August 20 by Cheng Xu, Nan Yan, Liming Chen, and M-Tahar Kechadi. The paper is about self-improving language models, but its most portable result is really about measurement.
A label such as learned looks as though it describes something that happened inside a model. In an evaluation system, however, the label is produced by more than the model. It also depends on how inference is served, how many answers are sampled, how answers are graded, and where the evaluator draws a line between states.
If those other parts can move the label, then a benchmark transition is not yet a model transition.
The batch was part of the instrument
Temperature-zero decoding sounds deterministic. At each step, choose the highest-probability token. Do that again from the same prompt and, in principle, you should get the same answer.
The principle is tidier than the server.
A NeurIPS 2025 study showed that greedy outputs can change when researchers change batch size, GPU count, or GPU type. The authors traced the effect to finite-precision arithmetic: changing how computations are grouped can change tiny rounding errors, and a tiny change early in a long reasoning trace can eventually produce a different answer.
A separate technical demonstration by Thinking Machines Lab made the systems effect unusually visible. The team ran the same prompt 1,000 times at temperature zero on Qwen3-235B. Ordinary inference kernels produced 80 distinct completions. Kernels engineered so that their numerical result did not depend on batch composition produced the same completion all 1,000 times.
The frozen-model experiment in Phantom Gains then tested batching inside its own evaluation. The researchers ran the same 200 problems twice. With as many as 64 requests in flight, 16 problem verdicts flipped between solved and unsolved. When the requests were strictly serialized, four flipped. Three-quarters of the apparent state changes vanished when the scheduling condition changed.
The last four did not vanish, and the authors do not pretend to know exactly why. They point to possible server-side batching or lower-level nondeterminism outside their control. That unresolved residue matters because it keeps the causal claim narrow. The experiment establishes that the inference system can change the observed verdict while the weights remain fixed; it does not establish one universal source for every flip.
This is the first change in perspective. A benchmark answer is not simply something a model emits. It is something a model-and-serving-system produces.
Zero is a sample result
Suppose the serving system becomes perfectly reproducible. There is still another way to manufacture a transition.
Ask a difficult question 128 times. The model never answers correctly. It is tempting to call the problem unreachable. After training, ask another 128 times. The model gets it right once. Now the problem appears to have crossed a boundary: formerly impossible, newly possible.
But zero successes in 128 attempts does not prove a zero probability of success. In the Phantom Gains setup, zero out of 128 still leaves an upper 95-percent confidence bound of about 2.3 percent on the underlying success probability. A rare answer can be real and simply fail to appear in the first sample.
This matters because some recent work deliberately studies the boundary of what a model can reach under repeated sampling. MATH-Beyond, published at ICLR 2026, filtered high-school math problems against common open models of up to 8 billion parameters under large sampling budgets, including pass@1024, to build a harder test of whether reinforcement learning can move beyond an observed baseline. That is a useful operational question. It is not the same as proving that the base model assigns exactly zero probability to a correct solution.
The frozen Qwen model makes the distinction hard to ignore. In one comparison, 25 AIME problems had produced no correct baseline answer in 128 samples. On the next evaluation, the unchanged model got seven of those problems right at least once. A rule that calls one later success an expansion therefore assigned the frozen model an expansion rate of 7 out of 25, or 28 percent.
The 28 percent is not a general false-expansion rate. It is one realization in one experiment. The more revealing result is that the problem survived replication inside the study. Across 110 frozen comparisons, even requiring at least two later successes left a measured background rate of 5.8 percent, with a reported 95-percent interval from 3.8 to 7.8 percent.
A separate ICLR 2026 Bayesian analysis of repeated-sampling evaluation reaches the same statistical caution from another direction: finite trial counts are evidence about an underlying success rate, and rankings can become unstable when those counts are treated as exact point estimates.
The important word in a phrase like the model could not solve this problem is often not model. It is could not, and the evaluation procedure has to say what that means.
Run the blank
Laboratory instruments have a familiar way to expose this kind of error. Measure the instrument when the thing you are trying to detect is absent.
For a transition ledger, the equivalent is simple: keep the model unchanged, push it through the same evaluation procedure, and count the transitions anyway.
Phantom Gains calls the resulting estimate a measured null. The terminology is statistical; the intuition is physical. If an unchanged system regularly produces three, five, or ten apparent changes, zero is the wrong baseline for interpreting a trained system.
The authors go further than pointing out the problem. They pool repeated evaluations of the untrained model and replace the one-success expansion rule with a per-problem statistical test. When they hold out frozen replicas and ask the new test to find changes, it finds none. That is what a no-change test ought to do.
Crucially, the correction does not make every training effect disappear. In the same study, distillation from a stronger external teacher improved 8 to 11 of 22 AIME problems that the base model reached only rarely, while three self-training variants improved 0 to 2. On the much smaller set of 10 problems that the pooled baseline never reached at all, the comparison was inconclusive. The strongest supported claim is therefore about rarely reached problems, not proof that the teacher pushed the model outside the base distribution’s support.
This is the point at which the measurement critique becomes useful rather than merely skeptical. A better instrument can preserve a real signal while stripping away a seductive false one.
It also defines a boundary for the prescription. The preprint reports that an ordinary analytic noise model agrees closely with the empirical frozen control for the solve-rate estimator the authors prefer. If an evaluator’s uncertainty is well understood, a statistical model may be enough. Re-running a frozen system is especially informative when several kinds of noise are compounded — serving behavior, sampling, grading, thresholds — and the convenient textbook distribution may not describe the full pipeline.
The rule is not “always duplicate this experiment.” It is “do not assume the no-change reading is zero.”
Real learning is the rival explanation, not the enemy
There is a larger debate behind these experiments. Does reinforcement learning mainly make answers already available to a base model easier to sample, or can it create genuinely new reachable behavior?
The evidence does not support a single answer for every training regime.
One NeurIPS 2025 study found that several reinforced models looked better than their bases at small sample counts, but that the bases often caught up or surpassed them when given many more attempts. That result supports a “sharpening” interpretation: training may move probability toward useful reasoning paths that were already present.
Yet ProRL, also published at NeurIPS 2025, reported a different result under prolonged reinforcement learning. Its trained models continued to beat their bases across large-sample evaluations, including problems the bases did not solve in the authors’ extensive sampling.
Those results are a reason to calibrate transition measurements, not a reason to dismiss transitions. If some training regimes really expand what a model can reach and others mainly rearrange probabilities, a noisy instrument makes that distinction harder to see.
The same caution applies to Phantom Gains itself. It is a fresh preprint centered on one Qwen3-8B backbone, math tasks, and a particular training-and-evaluation pipeline. Its frozen control is powerful evidence that the demonstrated false transitions are possible in that setup. It is not a prevalence study of all benchmarks, all inference stacks, or all self-improvement methods.
When a score becomes a state
The measurement problem becomes especially relevant when software turns an uncertain score into a durable label.
Some AI evaluations already rely on another language model to score an answer. An EMNLP 2025 paper on language models used as judges developed conformal prediction intervals around such scores, treating the judge’s output as uncertain rather than exact. If a system then thresholds an uncertain score into passed, failed, improved, or regressed, it has created the same kind of transition problem in a new setting.
That does not establish that production agent dashboards suffer the same false-transition rates as the math experiment. No evidence here supports that prevalence claim. It gives a testable prediction instead.
Freeze the system. Run the evaluator repeatedly under the same operational conditions. Ask how often the label changes anyway.
If the evaluator is deterministic end to end, the background transition rate may be zero. If it is stochastic, or if the serving system itself is not reproducible, the background may not be zero. Either result is useful because it tells you how much meaning a state change can carry.
This is the odd danger of granularity. “Accuracy rose from 65 to 68 percent” sounds statistical. “These fourteen capabilities were learned and these nine were forgotten” sounds diagnostic. The second statement is more specific, but specificity is not the same thing as resolution.
The frozen Qwen model never acquired six abilities. It never lost nine. Those events appeared only after the system converted two uncertain observations into a story about state.
Before reading that story, it helps to leave the model exactly where it is and ask the instrument to tell it anyway.
Sources
- Cheng Xu, Nan Yan, Liming Chen, and M-Tahar Kechadi, “Phantom Gains: Auditing Self-Improvement Against a Measured Null”, arXiv v1 preprint, submitted August 20, 2026.
- Jiayi Yuan et al., “Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference”, NeurIPS 2025.
- Horace He and Thinking Machines Lab, “Defeating Nondeterminism in LLM Inference”, September 10, 2025.
- Prasanna Mayilvahanan et al., “MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model”, ICLR 2026.
- Mohsen Hariri et al., “Don’t Pass@k: A Bayesian Framework for Large Language Model Evaluation”, ICLR 2026.
- Yang Yue et al., “Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?”, NeurIPS 2025.
- Mingjie Liu et al., “ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models”, NeurIPS 2025.
- Huanxin Sheng et al., “Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction”, EMNLP 2025.