Would a Different Step Have Changed the Outcome?
For long agent trajectories, the usefulness of a step-level score for training is separate from its fidelity as a measure of causal contribution; systems should validate a score against the job it will perform rather than infer causal meaning from local plausibility or downstream gains.
An AI agent is standing, in the loose sense available to text-based software, in a simulated bathroom. Its job is to complete a household task. It issues a command: take a cloth from a toilet. The environment replies that nothing happens.
The agent tries again.
Nothing happens.
It tries again, and then again.
What makes the episode useful is not that the agent is stuck. Agents get stuck all the time. It is that the model is almost perfectly confident about several of the repeated commands. In an August 20 preprint, researcher Haiyue Zhang reports confidence above 99.99 percent on three of them. A step score derived from the model’s own probabilities also treats the repetitions favorably.
Then Zhang asks a different question. She restores the simulation to the state before each command, substitutes other actions that the same model might plausibly have taken, and lets those alternatives run to the end. The repetitions stop looking equivalent. At two of those moments, keeping the repeated command produces a better measured outcome than the sampled alternatives. At the middle one, the alternatives do better.
The text of the command barely changed. Its place in the future did.
That small discrepancy is the most useful result in Credit Without Ground Truth, a new audit of step-level feedback for AI agents. The paper does not show that step scoring is useless. Other recent systems report the opposite: process feedback can improve training. What the replay experiment shows is more specific.
A score can be useful without telling you what caused the outcome.
What does a good step mean?
As AI agents take longer sequences of actions, researchers and developers increasingly want feedback somewhere between the prompt and the final result. If an agent searches five sites, edits a file, runs a test, fixes an error, and submits an answer, a single success-or-failure mark at the end throws away a great deal of information.
So systems score the steps.
But there is no single thing called a good step. A step can be factually correct. It can make progress. It can be necessary infrastructure. It can be redundant but harmless. It can look suspicious. It can make the final result more likely. Or it can be the decision such that, had something else happened there, the eventual outcome would probably have changed.
Those categories often travel together. That makes them easy to confuse.
The better process-reward work already avoids the simplest version of the mistake. A process-reward model published at The ACM Web Conference 2026 explicitly argues that an agent action does not have the clean notion of correctness that a line in a math proof might have; it instead evaluates ideas such as progress toward the goal. Another recent method, TRIAGE, sorts parts of a trajectory into roles such as decisive progress, useful exploration, no-progress infrastructure, and regression.
These are richer judgments than right or wrong. They are still judgments about what a step is doing in the observed trajectory.
A causal question asks for another object: if we changed this decision while holding a specified set of circumstances fixed, how would the outcome distribution change?
That distinction predates LLM agents. In reinforcement learning, counterfactual credit assignment has long been concerned with separating the influence of an action from what happened later or outside the agent’s control. Zhang’s contribution is to build an unusually direct audit for several signals now used around LLM agents: restore the environment, change the action, and actually run the alternatives.
The difference is easiest to see in a trajectory that can recover. An agent can make an ugly move and repair it three steps later. It can make an elegant move after the outcome has effectively become inevitable. A reviewer may reasonably dislike the first move and praise the second. Neither judgment, by itself, tells us how much replacing that move would have changed the final result.
Local quality and counterfactual leverage can coincide. They are not synonyms.
Rewind, replace, run
Zhang’s audit uses ALFWorld, a text environment in which an agent performs simulated household tasks. At each action turn in a collected trajectory, the experiment restores the prior state, re-executes the factual action, and samples four distinct alternatives supported by the same policy. It then runs the factual action and each alternative to a terminal outcome at least three times.
This matters because the counterfactual is executed rather than narrated. Another model is not asked to imagine what would have happened. The environment is actually rewound and rerun.
With a Qwen2.5-7B policy, the audit could construct the required set of alternatives for 1,768 of 2,034 intervened turns. Zhang then asked whether three ordinary signals—the model’s confidence, an outcome-conditioned probability score, and a much larger model acting as a judge—ranked steps the same way as the replay effect.
On the paper’s primary within-trajectory ranking test, they did not distinguish themselves from their matched shuffled controls. The outcome-conditioned score’s median rank correlation with replay effect was 0.019. The judge’s was 0.114. The judge was not completely uninformative: on one sign-based statistic it agreed 60.4 percent of the time. But it still failed the paper’s stronger preregistered test of concentrating high scores on the decisions with measured counterfactual leverage.
One reason appears in the model’s own probabilities. In the Qwen experiment, the outcome-conditioned score tracked the policy’s fluency far more strongly than it tracked the replay effect. Its median rank correlation with policy log-probability was about 0.75; after statistically accounting for that fluency signal, the remaining association with the replay increment was essentially zero. A second model family showed the same broad fluency-dominance pattern, though not every Qwen-side result transfers unchanged.
That makes the bathroom loop less silly than it first appears. Repetition is highly predictable. A system built partly from the model’s own probabilities can therefore reward a continuation because it looks very much like the continuation the model expected to produce. Predictability and influence are answering different questions.
The paper contains another number that requires equal care. Among the Qwen turns for which the replay counterfactual was defined, 30.5 percent had a stored non-zero effect. It is tempting to say that the other seven in ten decisions did not matter.
That is not what the experiment establishes.
Each action and alternative was sampled only a few times, outcomes are discrete, and the study defines a turn as pivotal whenever the stored replay effect is literally non-zero. For many turns recorded as zero, the sample size is too small to exclude modest hidden effects. Zhang is explicit that zero means indistinguishable at the experiment’s achieved resolution, not proof of no causal influence.
The useful finding is therefore a distribution, not a universal percentage: measurable outcome sensitivity is uneven, and the instrument can clearly distinguish some decisions from others. The exact prevalence of consequential steps is unresolved.
Even measurability depends on the policy. The experiment could not find the required set of policy-supported alternatives for 13.1 percent of intervened Qwen turns and 26.8 percent in the corrected Llama replay. A causal effect here is not a permanent property printed on the action. It depends on which alternatives are admitted, which policy continues the run, which environment is restored, and which outcome is measured.
A step matters relative to a counterfactual.
The strongest objection
There is an obvious response to all of this: perhaps causal fidelity is simply the wrong standard for a training signal.
That response has evidence behind it.
TRIAGE does not obtain its process labels by replaying every decision. A judge assigns semantic roles to trajectory segments, and those labels reshape the training signal. Yet its authors report higher success rates than outcome-only GRPO across ALFWorld, Search-QA, and WebShop for two policy models. On completed runs in two of those environments, the method also reduces the number of environment-facing turns.
Another preprint, CARL, starts from the idea that some states are more consequential than others but uses a cheap proxy during training: action entropy, roughly how uncertain the policy is among alternatives. The authors report that concentrating rollouts and updates on high-entropy states improves both performance and training efficiency relative to their baselines.
Other recent work points in the same direction. Fine-grained feedback can be useful because it reduces noise, shapes exploration, or gives the optimizer a denser signal. None of those benefits requires the score to be a perfect estimate of a decision’s marginal causal effect.
Zhang’s own training experiment makes this distinction harder to ignore. In a preregistered seven-arm comparison, no credit rule reliably beat the untrained policy. Some checkpoints initially appeared different, but the pattern was explained by training dose: sparse rules retained fewer examples and therefore received fewer optimizer steps. The experiment does not show that a more causally faithful score produces a better-trained agent.
So there are two claims a step score might earn.
The first is practical: using this score improves learning.
The second is interpretive: this score identifies the decisions that changed the outcome under a specified intervention.
Evidence for the first does not automatically establish the second. And failure on the second does not erase the first.
That is the stronger version of the argument because it permits both results to be true.
When the bad step really is the break
The distinction also shrinks under the right conditions.
Consider Who&When Pro, a benchmark for agent failure attribution. Its designers begin with a successful trajectory, exactly replay the successful prefix, and only then inject a failure. Using that controlled pipeline, they construct 12,326 failed trajectories with known labels.
The setup changes the question. Instead of observing an organic run containing recovery, redundancy, and several possible mistakes, the experiment deliberately creates a point after a verified successful prefix where the trajectory is made to fail. Error localization is therefore much closer to locating the break in the successful path.
That is not a defect; it is a boundary condition. If a workflow has one forced failure after a known-good prefix and no meaningful recovery path, “where did it go wrong?” can be a very useful approximation. In a branching workflow with retries, repairs, delegated work, and interacting errors, the same shortcut is less secure.
There is a second boundary: sometimes counterfactual replay is unusually feasible. Causal Agent Replay formalizes an agent run as a causal system, changes a step, and re-executes the continuation; its author validates the estimators on synthetic cases where the causal structure was planted in advance.
And in cooperative multi-agent LLM systems, C3 exploits a special property of text-only interaction histories. At a decision point, it can restore the complete observable history, sample alternative actions under a frozen behavior policy, and compute per-decision advantages. In its current revision, C3 reports better results than its baselines across six benchmarks spanning mathematical reasoning and code generation.
Together these cases show why “use causal replay” is no more universal than “trust the process score.” Replay works best when the relevant state can actually be restored and repeated. Zhang’s ALFWorld audit spends multiple rollouts on every factual and alternative action. C3 depends on interaction history that is fully observable in text. A production system may depend on APIs, people, changing databases, physical devices, or one-time external events. Sometimes the world will not sit still for the experiment.
The right instrument depends on the job and on what the system permits you to measure.
Three jobs hiding behind one score
That leaves a practical problem for teams building agents. Step-level feedback is often presented as one feature: score the trace. But the same score can be asked to perform at least three different jobs.
For training, the important question is whether the signal improves the resulting policy under a fair comparison. If two methods expose the model to different amounts of training, that difference must be controlled before crediting the score itself.
For diagnosis, the question is whether the explanation survives plausible interventions. A score that accurately flags regression may still be a poor answer to “which decision changed the outcome?” if the system recovered from that regression every time.
For review allocation, the requirement is stricter still. If scarce human or model attention is supposed to be routed toward places where changing a decision has high expected value, then the routing signal should be validated for that purpose. Suspiciousness, uncertainty, progress, and causal leverage are all plausible triage signals. None gets the causal interpretation for free.
This is less convenient than having one universal definition of a good step. It is also more useful. It turns an argument about which score is best into a prior question: best for what?
Return to the simulated bathroom. The agent takes the cloth and tries again. From inside the model, the repeated commands are nearly certain. From the replay experiment, their measured effects differ because each sits at a different point in a trajectory whose future is still changing.
The model had no trouble recognizing what it was likely to do next.
Finding the place where doing something else would change the future was a different task.
Sources
- Haiyue Zhang, “Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay”, arXiv preprint, submitted August 20, 2026; author lists the manuscript as under review.
- Zhiheng Xi et al., “AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress”, The ACM Web Conference 2026, pp. 4184–4195. DOI: 10.1145/3774904.3792551.
- Yuanda Xu et al., “TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning”, arXiv preprint, submitted June 30, 2026.
- Leyang Shen et al., “CARL: Criticality-Aware Agentic Reinforcement Learning”, arXiv preprint, submitted December 4, 2025; revised May 11, 2026.
- Jiale Liu et al., “Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?”, arXiv preprint, submitted July 10, 2026.
- Jaineet Shah, “Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures”, arXiv preprint, submitted June 6, 2026.
- Yanjun Chen et al., “Exact Is Easier: Credit Assignment for Cooperative LLM Agents”, arXiv preprint, submitted March 6, 2026; revised May 8, 2026.
- Thomas Mesnard et al., “Counterfactual Credit Assignment in Model-Free Reinforcement Learning”, Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021.