Skip to content

Counting the Same Evidence Twice

Related:Cards ↗

When an agent workflow treats multiple memories, reviewers, or model outputs as corroboration, raw agreement can overstate the evidence if those records repeat the same query-relevant support; provenance can reveal shared lineage, but it does not by itself establish evidential independence.

The researchers changed the evidence without changing the truth.

They started with three benchmarks for long-term AI memory. Then they added records. Some were genuinely different pieces of evidence. Others were paraphrases and summaries derived from material already in the memory pool. The correct answers stayed the same.

A simple voting system did not always stay put.

In a new preprint called Beyond Memory Majority, the researchers measured how often an answer changed after redundant, correlated memories were added. With one model as the underlying engine, majority voting changed its prediction on 44.8 to 49.3 percent of the tested cases, depending on the benchmark. With a second model, the range was 46.5 to 51.0 percent. The authors’ correlation-aware method reduced those ranges to 7.8–10.3 percent and 8.6–11.2 percent, respectively.

Nothing about those numbers tells us how often this happens in the wild. The paper appeared on arXiv on August 20, 2026, and the benchmark variants were deliberately built to contain duplicated descendants. There is no independent replication yet, and the experiment does not measure the prevalence of false majorities in deployed agent systems.

But it isolates an unusually clean failure. The system can receive more agreement without receiving another reason to be right.

Why a rewrite gets another vote

Suppose an agent reads a report and stores a conclusion. Another agent summarizes the conclusion in a project note. A third uses the note in a plan. A fourth stores a memory about the plan. Later, a verifier retrieves all four records.

Four pieces of text are now on the table. They may still trace back to one relevant observation.

That is not the same problem as having four bad voters. Nor is it simply the familiar fact that majorities sometimes make mistakes. The unusual feature is that the workflow itself has multiplied the apparent evidence. A rewrite becomes a new record, and a new record can become another vote.

This bookkeeping error predates AI by a long time. Cochrane’s handbook for systematic reviews warns researchers not to count multiple reports of the same study as multiple studies. One clinical trial can produce a paper, an abstract, a follow-up report, and other publications. The documents multiply; the experiment does not. Cochrane therefore tells reviewers to link the reports and keep the study as the unit of interest.

An agent memory is not a clinical trial. The useful parallel is narrower: the number of documents is not necessarily the number of observations.

Agent workflows make document multiplication cheap. Summaries, extractions, plans, reports, and stored memories can all descend from material that came earlier. If ancestry disappears along the way, later software sees several polished artifacts and has little reason to know that they are relatives.

A family tree does not solve the whole problem

The obvious fix is to preserve provenance: keep track of where each memory came from. Then, perhaps, the verifier could trace five agreeing records back to one root and count them once.

The new memory paper argues that this rule is too crude.

One source can contain several different pieces of useful evidence. A quarterly report might contain a revenue figure and a description of a supply interruption. If a question depends on both, “one source, one vote” would throw information away. The reverse is possible too. Two records with different origins can still carry the same bad premise or reflect systems that tend to fail on the same cases.

So the thing that matters is not a source ID. It is what a piece of evidence adds for the question being decided. The paper’s method tries to estimate those query-specific groups of support, using provenance as one clue rather than as a hard definition of independence.

That makes the image of one witness repeated ten times useful, but incomplete. Sometimes ten records really do descend from one observation. Sometimes one source contains several nonredundant facts. Sometimes ten nominally separate systems share a blind spot.

Provenance can show family resemblance. It cannot certify independence.

Nine models, roughly two effective judges

A separate study makes that last point without using a memory family tree at all.

In May, researcher Guneet Kohli tested nine frontier language models from seven model families as judges on three natural-language-inference datasets. The models were nominally diverse. Their errors were not independent.

Using the panel’s observed error correlations, Kohli calculated an effective sample size of roughly 2.2 to 2.5 judges across the three datasets. The majority panel performed 7.6 to 22 percentage points worse than an idealized panel whose votes were independent. On one dataset, the nine-model majority edged the best single judge by 0.2 percentage points. On the other two, the best single judge did better.

The estimate is not a law of model panels. The study uses particular classification tasks, particular models, and gold labels based on 100 human annotations per item. Its own statistical interval applies to those judges and those items, not to some universal nine-becomes-two conversion rate.

What it supplies is an independent demonstration of the mechanism’s second half. Dependence need not be visible in a lineage graph. Nominally separate systems can still make the same mistakes on the same examples.

Now the problem is larger than duplicate memories. A workflow can have many reviewers and still gain little assurance from the next one if the new reviewer brings the same information and the same failure pattern.

Maybe majority voting is simply the wrong tool

There is a strong rival explanation for all of this: perhaps the real problem is not evidential lineage. Perhaps majority voting is just a crude aggregation rule.

There is good evidence for that criticism. A paper accepted at ICML 2026 develops aggregation methods that use information about model quality and cross-model correlation rather than treating every answer equally. Across synthetic data, standard language-model benchmarks, and a healthcare application, the proposed methods outperform ordinary majority voting. Another recent preprint, AgentAuditor, reports gains from examining where agents’ reasoning diverges instead of merely counting final answers.

Those results cut against any simple story in which voting itself is doomed. Better aggregation can recover useful information from a panel that raw majority voting mishandles.

But they also sharpen the more defensible lesson. Once errors or evidence are dependent, headcount is no longer enough information to interpret the headcount. One system may address that with provenance. Another may estimate correlation. Another may inspect the underlying reasoning or check the claim against external state. The common move is to ask what each additional agreement contributes instead of assuming that one more answer equals one more unit of corroboration.

That is a much narrower claim than “more agents do not help.” More agents can help a great deal. The question is why.

The audit trail has to survive the rewrite

Provenance remains useful even though it cannot answer that question by itself.

A memory system that preserves ancestry can distinguish three records from distinct source episodes from three summaries derived from one episode. Without that history, the derivation has been flattened into prose.

At least some production systems already treat this lineage as an engineering primitive. Zep, a commercial agent-memory provider, described in July how its system links derived facts back to the source episodes that produced them. The company says those associations support source-scoped retrieval, access control, debugging, and deletion propagation. That is a vendor’s account of its own implementation, not evidence that false majorities are common. It shows something more modest: keeping the family tree is technically possible.

The harder judgment comes afterward. Two memories with the same parent may carry different relevant facts. Two memories with different parents may repeat the same claim. Two models with different names may share an error pattern. A graph can tell the verifier where the material came from; it cannot, on its own, say how much new evidence the material contains.

For low-stakes work, that distinction may not justify extra machinery. And the new memory study gives us no denominator for how often ordinary agent workflows suffer from the failure it constructs. The important boundary is a consequential decision in which several memories, reviewers, or model outputs are being presented as corroboration.

At that moment, “How many agree?” is incomplete.

The more useful question is: What did the last agreement add?

Return to the opening experiment. The researchers did not change the answer. They added descendants—new pieces of text whose relevant support was inherited from evidence already present. A majority system often changed its decision anyway.

The software had not learned another fact. It had learned to count a rewrite twice.

Sources