What Gets Lost When Agents Pass a Claim Along
Agent handoffs are most worth checking when they erase provenance, compress evidence, or create downstream fan-out; preserving claim lineage can make selected intermediate checks useful, but current evidence does not support verifying every handoff or replacing terminal review.
A recent financial-question-answering system checks the same work twice.
CLAIR-Fin, an August 2026 arXiv preprint, breaks questions about long financial reports into small claims and collects evidence for each one. Before a claim moves from drafting into adversarial review, the system checks whether the draft is still supported by that evidence. Later, near the end of the pipeline, it checks again.
At first this looks like expensive caution. If the final audit works, why bother checking a claim in the middle?
The authors’ ablations make the duplication harder to dismiss. On a 500-question benchmark built from Bangladesh Bank annual reports, the full system scores 0.889 on an automated RAGAS semantic-faithfulness measure. Remove the intermediate handoff check and the score falls to 0.857. Remove the terminal audit instead and it falls to 0.845. On exact correctness, however, the ranking flips: the full system scores 0.592, compared with 0.548 without the handoff check and 0.561 without the terminal audit.
There is no clean winner. Different metrics favor different checks. The more interesting result is that the intermediate check changes the outcome at all while a final audit remains in place.
Why?
A claim can lose its evidence without becoming false
Verification at more than one stage is not new. Chain-of-Verification, published in Findings of ACL 2024, has a model draft an answer, generate questions that could expose mistakes, answer those questions independently, and then revise. Claim-level checking is also well established enough that the novelty here cannot simply be “look at smaller pieces.”
The more useful question is what changes when one component’s output becomes another component’s input.
A research agent may begin with a table, a source passage, and a caveat. A writing agent may receive one sentence. A planner may produce several pages of reasoning and hand an executor a short instruction. A coding agent may turn a local assumption into an interface that other agents begin implementing.
Each transition is a form of compression. Compression is useful; without it, the next component would have to repeat the entire investigation. But compression can discard the information that says why a conclusion deserves to be trusted.
A 2026 arXiv preprint called ProvenanceGuard isolates one version of this problem. Its verifier keeps stable identities for tools and sources, then asks two separate questions: is a claim supported, and is it supported by the source to which it is attributed?
The distinction sounds fussy until the source is wrong but the sentence is right. In 50 controlled source-conflation probes, the system detected every injected attribution swap. Yet on a harder multi-source benchmark, exact source-and-relation accuracy fell to 0.229. Preserving a source ID helps; figuring out which of several similar sources actually owns a claim can still be difficult.
Imagine a research agent reports a statistic accurately but assigns it to the wrong study. A writing agent copies the number. A final reviewer searches for the number, finds it elsewhere, and decides the sentence is supported. The fact survived. Its evidentiary lineage did not.
That difference matters when source identity carries meaning: a regulator rather than a vendor, an original experiment rather than a summary, a production log rather than a synthetic benchmark. A final answer can be factually plausible while obscuring how it became plausible.
Fan-out changes the cost of a mistake
Losing provenance is one problem. Reusing a conclusion creates another.
The arXiv preprint From Spark to Fire models multi-agent collaboration as a graph of message dependencies. Its experiments are deliberately artificial: the researchers inject a single false claim into short, controlled agent networks and watch what happens. In five of six tested framework settings, some configurations reached 100 percent final infection. Their genealogy-based controls reduced propagation sharply, but stronger controls also consumed more tokens and added latency.
This is not evidence that ordinary software-agent teams routinely collapse into false consensus. The study uses synthetic error seeds, short horizons, and a simplified evaluation of whether an agent adopted the false claim. What it does show is that an error can acquire a different practical importance after other work begins to depend on it.
A wrong value in one report is one defect. The same wrong value copied into shared configuration is a different class of problem. The number did not become more false. It acquired fan-out.
That is why “check at the end” and “check before reuse” are not always interchangeable. A final reviewer may still catch the original mistake, but by then the mistake can have generated more work: summaries, patches, plans, decisions, or further claims that have to be unwound separately.
The handoff itself is not magical. The important transition is the one that turns a local conclusion into shared state.
Sometimes the next agent fixes the problem
There is a strong rival explanation to the whole handoff story: perhaps another agent is not mainly a source of propagation. Perhaps it is another chance to correct the work.
A separate 2026 arXiv preprint, Hallucination Cascade, ran 500 cascade experiments across ten knowledge domains. In three-agent chains, the paper’s normalized hallucination score fell from 0.422 at the first agent to 0.272 at the third. By that measure, later stages attenuated hallucination rather than amplifying it.
But factual accuracy also fell, from 0.789 to 0.769. The later agents produced answers that looked better on one error measure while preserving slightly less factual content on another.
That result is useful precisely because it refuses the tidy story. A downstream agent can correct, omit, distort, compress, or ratify what it receives. “Another agent touched it” does not tell us which of those things happened.
It also makes a universal checkpoint rule hard to defend. CLAIR-Fin shows an intermediate grounding check adding protection in one financial-QA architecture. From Spark to Fire shows how synthetic errors can spread through dependency structure. ProvenanceGuard shows that support and source ownership are different questions. Hallucination Cascade shows that later agents can also reduce one class of error. These are benchmark and controlled results, not evidence that every production handoff deserves a gate.
The missing experiment is straightforward to describe and harder to run. Hold the models, prompts, tools, and verification budget constant. Compare final audit only, claim-level final audit, selected handoff checks plus a final audit, and checks at every handoff. Then measure not only final correctness, but propagation depth, rework, latency, tokens, and false blocks. Until someone does that, “verify every handoff” is a slogan, not a finding.
The risky handoff is the lossy one
What the current evidence does support is a narrower design question: what will be harder to recover after this transition?
A handoff deserves more scrutiny when it compresses or translates evidence, when source identity or uncertainty is likely to disappear, when many downstream tasks will reuse the result, or when the next action will be expensive to undo. Those conditions are an inference from the studies above, not a validated formula. But they explain why some transitions are more consequential than others.
A research-to-writing handoff may need the exact source passage and attribution to travel with the sentence. A planner-to-executor handoff may need the constraints that made the plan safe. A code-producing agent may need tests or durable context attached to the regions another agent is about to modify.
The common feature is not the presence of another agent. It is the conversion of someone else’s conclusion into a premise.
This also separates persistence from verification. One example from our own work is git-span, a Rust CLI that stores relationships between exact code anchors as ordinary tracked files and can report when an anchor has drifted. It does not prove that a relationship is correct. It preserves the relationship so a later agent or person can still inspect it.
That is a smaller claim than reliability, and a more useful one. A perfectly preserved mistake is still a mistake. But a mistake with a visible origin and visible dependents is easier to challenge than one that has dissolved into context.
What the polished answer leaves behind
CLAIR-Fin has an object that never appears in its final prose: a claim ledger. The ledger persists through the workflow, carrying evidence, provenance, and verification state while the wording around it changes.
That may be the most important part of the system—not because the ledger makes the answer true, and not because the intermediate check beats the final one. It preserves enough of the investigation that a later component can still ask what kind of object it has received.
The final answer is the compressed product. The ledger remembers what the compression removed.
As more work passes from agent to agent, the handoff worth designing for is the moment a sentence stops asking to be proved and starts being used. The safest thing to pass is not just the sentence. It is the sentence, its evidence, and enough lineage for the next agent to know that it is still a claim.
Sources
- Fatema Tuj Johora Faria et al., “CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA”, arXiv preprint, August 13, 2026.
- Yizhe Xie et al., “From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration”, arXiv preprint, submitted March 4, 2026; revised May 11, 2026.
- Ander Alvarez et al., “ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents”, arXiv preprint, submitted June 16, 2026; revised July 26, 2026.
- Saeid Jamshidi et al., “Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems”, arXiv preprint, June 6, 2026.
- Sen Xu et al., “Claim-Level Reliability Assessment for Efficient Test-Time Reasoning”, arXiv preprint, August 12, 2026.
- Shehzaad Dhuliawala et al., “Chain-of-Verification Reduces Hallucination in Large Language Models”, Findings of the Association for Computational Linguistics: ACL 2024, August 2024.
- Goodfoot Media LLC, git-span repository and README, accessed August 18, 2026.