Skip to content

Same Transcript, Different Continuation

For the open LLM serving systems studied so far, application-layer provenance can omit serving state that changes an agent trajectory; exact replay requires controlling that state, while many engineering workflows may be better served by artifact-level rather than token-level equivalence.

Alexander Boesgaard Lorup started with a situation that should have been boring: the same model had reached the same sequence of tokens twice.

In one case, the model still had the key-value cache created while those tokens were originally generated. In the other, the cache had been reconstructed by feeding the identical token sequence through the model again. To anyone reading the transcript, the two states looked the same.

They did not behave the same.

In a matched 200-item experiment, the two continuations differed on 166 items under BF16 arithmetic. Twenty items crossed a correctness boundary. Exact repetitions of either computational path remained stable.

Then Lorup moved the hidden state itself. He transplanted the full cache from one path into the other. In 24 of 24 selected divergent cases, and later in 43 of 43 cases chosen without looking at the outcome, the continuation followed the cache donor.

The visible conversation had not changed. The computational history had.

That is a useful complication for a field becoming increasingly serious about agent provenance. We can save the prompt, the transcript, the tool calls, the repository, the model identifier, and a seed. We can still fail to save everything that made this exact run happen.

Two repeatable answers to the same visible request

A September 4 arXiv v1 preprint by independent researcher Aditi Patodiya tests the same boundary in a multi-turn serving system. The paper, submitted to IEEE for possible publication, is only days old, and several of its measurements required correction after collection. Its strongest controls are nevertheless unusually revealing.

Patodiya ran an 80-episode tool-use workload while fixing the model, decoding settings, seed, request order, batch size, hardware, and serving-engine version. Requests were issued serially. With caching disabled, repeated executions were bit-identical across every tested configuration: 0 of 800 episodes diverged.

When cached and recomputed paths were compared under llama.cpp with Qwen2.5-7B, the trajectories separated. In the original grid, they differed on 36.2 percent of episodes at F16 and 75.0 percent at Q4_K_M. After the experiment disabled a separate host-memory prompt cache that had confounded the first measurement, the cross-path difference remained: 40.0 percent at F16, 81.2 percent at Q4_K_M, and 77.5 percent at Q3_K_M.

The low-bit weights amplified the effect, but they were not its sole cause. Divergence was already present at F16, and the two lowest-precision cells did not form a clean monotonic progression.

The most important control was simpler. Patodiya restored the relevant server state and repeated each path. The recomputed path reproduced 40 of 40 times. The cached path reproduced 40 of 40 times. Yet the two paths still disagreed on 14 of those 40 items.

Both answers could be deterministic.

The request alone did not tell you which deterministic system state would answer it.

The cache was not one thing

That distinction matters because “the cache” is too neat a culprit.

Patodiya’s first llama.cpp experiment contained two different kinds of reuse. One was the attention-state reuse the study meant to compare. The other was a host-memory prompt cache that persisted conversation states across requests and selected among them by prefix similarity.

With that second cache active, one repeated-run comparison diverged on 38.8 percent of episodes. With it disabled, the same comparison fell to 1.2 percent. Intervening traffic strongly changed later results when the host-memory layer was active and had almost no effect when it was not.

So there are at least two phenomena to keep separate. First, a retained cache and a recomputed cache for the same visible tokens can be numerically different. Second, a long-lived server can carry reuse state from earlier requests that changes which internal path a later request takes.

The broader category is not “bad KV caching.” It is serving state that the application request does not name.

Independent work supports that narrower mechanism. Ranjith Chodavarapu and Lei Xu found systematic divergence between cached and recomputed FP16 inference across three open-weight models under greedy decoding; their FP32 control eliminated the observed token flips. Lorup’s transplantation experiment goes further by moving the candidate causal state and watching the continuation move with it.

None of these studies says that every cache implementation has this property or that a token prefix must always map to multiple useful outcomes. They show that, in the tested systems, the visible prefix can fail to identify a unique decoder state.

A tiny numerical difference can cross an interface

Why should two computational routes through the same tokens disagree at all?

Inference is floating-point arithmetic. Reusing stored attention state and recomputing it can change the order of reductions, accumulations, and kernel operations. Floating-point addition is not perfectly associative, so mathematically equivalent constructions need not be bit-identical.

Most such differences are tiny. A decoder only needs one of them to land near a decision boundary. If two candidate tokens are close enough, a small shift can change which one wins.

In ordinary prose generation, that might replace one acceptable word with another. In an agent loop, a token can sit inside a function name, an argument, or a command. A different tool call produces a different observation. The observation becomes part of the next prompt. The numerical difference has escaped the model and entered the environment.

This is the mechanism that makes multi-turn agents interesting here. It is also where the evidence becomes thinner.

Patodiya’s agent benchmark had low task success—between 1.2 and 18.8 percent in the reported cells—so it cannot show that hidden serving state frequently changes useful work. Its strongest outcome analysis comes from a separate math workload. Across 1,300 distinct items, the paper found 20 cache-path correctness flips: 13 favored the cached path and seven favored recomputation. It did not detect an aggregate accuracy penalty from caching.

A different study points the other way. Chodavarapu and Xu found cache-on accuracy higher in eight of nine tested conditions. The robust result across these papers is path divergence, not a universal direction of quality change.

A different token is therefore not automatically a defect. For a coding agent, the consequential question is whether the fork changes a tool call, a patch, a test result, or something a reviewer would care about.

The backend was never transparent

Caching is only one reason the application-layer record can be incomplete.

David Pape, Jonathan Evertz, and Lea Schönherr held model weights, decoding settings, and hardware constant while changing the inference engine. Across five backends, their 2026 study found benchmark-score shifts as large as 16.6 percentage points. They traced the differences to a bundle of serving choices including kernels, CUDA graphs, prefix caching, and engine-specific defaults. Their literature survey also found inference-stack details rarely reported.

That makes “the run” a boundary-setting problem.

A Git commit is obviously part of a coding-agent run. So are the tool results that caused the agent to edit one file rather than another. The model version and prompt belong there too when reproducibility matters. The inference engine, numerical precision, cache policy, and server lifetime have historically looked more like implementation details.

The serving studies show that some of those details can be causal.

This does not mean we should preserve every byte of server memory. A provenance system is useful because it selects the state worth keeping. Recording a full attention cache for every request would be expensive and, for many purposes, pointless. A known cold start, a deterministic execution mode, or a server-reset policy may be a better control than trying to archive the hidden state itself.

And there is an even more important escape: we can stop pretending that every workflow needs token-identical replay.

There is more than one kind of “same run”

For debugging a rare tool-selection failure, exact trajectory replay can matter. A benchmark author may need server state controlled tightly enough that two experimental conditions are not accidentally separated by warm-state history. A forensic investigation may need to know which engine and reuse policy produced an action.

A product team may want something looser. Two coding-agent runs can differ token by token and still produce the same Git tree. They can produce different trees that pass the same tests and preserve the same invariants. They can even pass the same tests while differing in a way a reviewer later decides matters.

Those are different equivalence targets.

This is where the strongest rival to the hidden-state story becomes useful. Deterministic inference is not impossible. Thinking Machines Lab has demonstrated batch-invariant operators that remove an important source of serving nondeterminism. SGLang has implemented deterministic inference while retaining some serving optimizations, although its published performance tests disabled radix caching for some kernel backends and reported substantial latency penalties. LLM-42 pursues a different route using verified speculation and rollback.

These systems make the lesson less mystical. LLMs are not doomed to irreproducibility because they are probabilistic objects. Repeatability is partly an engineered property of the complete inference system, purchased with particular controls and costs.

Caching has equally concrete benefits. A 2026 evaluation of more than 500 long-horizon agent sessions found that, using the best reported cache mode for each tested model, provider-level prompt caching reduced API costs by about 41 to 80 percent and improved time to first token by about 6 to 31 percent. Individual strategies varied more widely, and one full-context GPT-4o condition was slower. Those hosted prompt caches are not the same mechanism as Patodiya’s local KV-cache experiments, but the economic pressure toward reuse is obvious.

The engineering question is not “cache or correctness.” It is which kind of repeatability is valuable enough to pay for.

What the transcript can and cannot promise

Lorup’s transplantation experiment gives the problem a physical form. The transcript told us which tokens the model had seen. It did not tell us which internal state represented those tokens. Moving that hidden state moved the continuation.

Patodiya’s reset experiment supplies the other half. Once the hidden state was controlled, each of two disagreeing paths could become repeatable.

For self-hosted systems that require exact replay, a practical run manifest may therefore need the engine and version, weight and cache precision, deterministic-mode settings, cache policy, server-reset discipline, and enough telemetry to know whether prior state was reused. For hosted frontier APIs, the studies here do not establish what can be observed or controlled; provider behavior cannot be inferred from llama.cpp or vLLM.

For many coding-agent systems, the better answer may be to record the application state rigorously and define replay at the level of the durable artifact: same tree, same tests, same externally visible behavior. If hidden serving state changes the transcript but not the work, exact numerical provenance may be an expensive answer to the wrong question.

But when the exact path does matter, a transcript is not a complete causal record merely because it contains every word.

It records what the model saw.

It may not record how the server arrived there.

Sources