When an AI Teammate Leaves, Whose Memory Goes Stale?
In LLM teams that explicitly persist partner-specific notes, changing a teammate can invalidate state held by the incumbent, so task performance can recover before coordination efficiency; the effect is bounded, architecture-dependent, and not inevitable.
A recent experiment replaced one member of an AI team and then counted who had to compensate.
The replacement looked unusually safe. It used the same model as the departing agent, had the same role prompt, and had accumulated the same amount of experience. It was not a novice. It had simply learned to work with a different teammate.
In the harder version of a collaborative cooking task, that was enough to raise the team’s communication cost per completed sub-task by about 51 percent relative to a placebo roster change. In two-player Hanabi, the increase was 63 percent. The task scores fell too, but much less than the coordination-efficiency measure. The September 4 arXiv preprint was designed to separate those two outcomes.
Then the researchers asked where the extra communication came from.
When they replaced the agent that normally set the plan in the cooking task, roughly 70 percent of the added messages came from the agent that had not been replaced.
The worker who stayed was the one re-explaining, reconfirming, reporting status, and reissuing requests.
Nothing about that agent’s role had changed. Its teammate had.
That was enough to make some of what it remembered wrong.
A fact with another agent in the key
The study, by Jianxin Gao and four colleagues, formed independent two-agent teams from the same frozen base model. Each team worked together for ten formation episodes. After every episode, each agent rewrote a private 200-token notebook that the system would feed back into future sessions.
The notebook had two headings. One was for the task: recipes, affordances, recurring failures. The other was for the partner: what this teammate tends to do without being asked, what must be stated explicitly, which handoff phrases work, which signals have acquired a shared meaning.
The researchers then traded role-matched agents between teams. A placebo condition performed the same roster-change announcement and context reset but put the original agent back in the seat. That matters because even announcing a personnel change had a measurable cost of its own.
The swap created a useful distinction. Some learned information traveled well. An experienced stranger was generally much better than an empty-notebook novice. But some information was tied to the pairing.
Consider a note like: the responder reports ingredient status without prompting. It is stored in one agent’s notebook, but its truth depends on another agent. Replace the responder and the sentence may remain perfectly readable while its referent has changed underneath it.
This is a different kind of persistence from a fact about a repository or API. The bytes live in one place; the validity condition lives across two places.
Call it partner-conditioned state.
The study did not have to infer that category only from aggregate scores. Because all adaptation lived in text rather than model weights, the researchers could delete just the partner-note section. In the more coordination-heavy settings, clearing an arriving agent’s notes about its former teammate improved the swap condition. An agent sometimes did better after forgetting the wrong partner.
That result is suggestive, but not clean causal proof of semantic interference. Deleting the notes also shortens the prompt. Independent work has found that longer context can itself hurt model performance even when the needed information remains retrievable. Gao and colleagues did not run a same-length neutral-note control. Their notebook template also explicitly asks agents to maintain a partner section, which may encourage more relationship-specific state than a production system would create on its own.
The statistical base is small as well. Each main cohort contains eight teams, but the role-matched swaps form only four independent team pairs. The authors bootstrap those four pairs for the central swap contrasts. And the two cooking conditions differ in overall task complexity as well as required coordination, so their ordering cannot be read as a clean causal estimate of interdependence.
The paper establishes a possibility under a deliberately persistent memory architecture. It does not establish how common that possibility is in deployed agent systems.
The old idea hiding inside the new one
Human organizations have a name for a related phenomenon.
Research on transactive memory studies how groups learn not only facts but also who knows what and how work moves among particular people. A 2005 study of hospital joint-replacement teams found that experience working together contributed to performance separately from individual and organizational experience. A later study of software-development teams found that team familiarity had a significant positive relationship with performance, while conventional measures of individual experience were not consistently predictive.
That does not mean a language model with a 200-token notebook has a human social relationship. The useful similarity is narrower: repeated work can make information about this collaborator economically valuable, which means a membership change can invalidate information without changing the underlying task.
The human literature also supplies an important corrective. A review of 133 empirical studies of team membership change since 1948 found that membership changes often disrupt coordination and shared cognition at first, but can improve performance after teams adapt. New members can add knowledge and capability that ultimately outweigh the transition cost.
So the engineering question is not whether teams should remain frozen.
It is whether we can tell the difference between the value of the new member and the cost of changing the relationship around them.
The cost that a success dashboard misses
That distinction becomes visible in the recovery curves.
After the swap, task score returned to within one percent of the placebo baseline by episode two in the easier cooking condition, episode five in the harder one, and episode six in Hanabi. Coordination cost recovered more slowly. In Hanabi, teams were still paying about a 7 percent premium at episode ten; neither of the two more coupled settings had returned within one percent of baseline coordination cost by then.
A success-only dashboard would therefore pronounce these teams recovered before their old efficiency had returned.
The paper can also say what some of the extra traffic was doing. In the easier cooking condition, one third of added messages were requests repeated after no response and one quarter were clarifications. In the harder condition, the largest category—38 percent—was correction after a handoff had already failed.
That is why the 51 and 63 percent headline numbers need care. They are not raw message increases. The metric divides communication by progress, so both numerator and denominator move. Raw communication rose 38 percent in the harder cooking condition and 42 percent in Hanabi, while progress also declined. The ratio captures what an operator ultimately pays per unit of completed work.
This is one advantage of instrumented agent teams: some coordination costs that are fuzzy in a human organization are directly meterable. Messages, hints, tokens, and latency can be counted.
But precise measurement does not help if the dashboard watches the wrong thing.
Some agents never learn the private dialect
The strongest rival evidence comes from systems where the effect does not appear.
In the 2025 LLM-Coordination benchmark, language-model agents were relatively robust to unfamiliar partners in zero-shot coordination. In Hanabi, an LLM paired with a partner it had not practiced with did not show the severe cross-play collapse seen in some reinforcement-learning agents.
A 2026 repeated-reference-game study offers an even sharper contrast. Human pairs gradually compressed their descriptions into efficient, partner-specific shorthand. The multimodal LLM pairs did not. Their descriptions stayed verbose, and their label overlap was statistically indistinguishable between real partners and pseudo-pairs with no shared history.
Those agents coordinated without forming the compact private conventions humans did.
That means partner-conditioned state is not an intrinsic property of two LLMs talking to each other. Gao’s result appears in a more specific regime: repeated interaction, persistent cross-episode memory, an explicit place to write partner expectations, and tasks where coordination can reward learning what the other participant tends to do.
The paper’s own ablations point in the same direction. In one high-coupling setting, greedy decoding reduced both the divergence between independently formed teams and the normalized swap penalty. Extending formation from ten to twenty episodes increased the measured partner-specific residue, though the authors call the trend preliminary. More opportunity to form conventions was associated with a more expensive swap—up to the point where very high sampling temperature also damaged task performance and made the ratio unstable.
A fixed communication protocol is therefore a serious rival to better memory management. If agents can be made interchangeable by standardizing what must be said and when, then the cheapest solution may be to prevent private conventions from becoming necessary rather than to preserve and transfer them forever. The paper does not test that intervention; it identifies it as future work.
Memory needs more than a timestamp
Most persistent-memory systems are organized around accumulation. Useful facts survive. Old context can be summarized. Important lessons move forward.
The swap experiment suggests another requirement: memory needs an invalidation key.
A repository fact may be keyed to a commit or file version. An API fact may be keyed to an interface version. A partner expectation may be keyed to an agent identity—or, more precisely, to a relationship between two participants.
Those facts can all be written in the same Markdown document and still have different lifetimes.
This is where the worker who stayed becomes the important figure in the experiment. On a roster change, it is easy to focus on what the newcomer lacks: the plan, the codebase history, the local conventions. But the incumbent can have the opposite problem. It may possess context that is no longer true.
Replacing a teammate can therefore create two onboarding tasks at once: teach the newcomer what remains valid, and tell the survivor what just expired.
The paper’s most revealing result is not that an unfamiliar agent talked more. It is that, when the planning seat changed, most of the added talking came from the seat that did not.
The system had replaced one agent.
The stale state showed up in the other one.
Sources
- Jianxin Gao, Tianyi Yu, Linna Deng, Runze Li, and Zining Wang, “Testing Interchangeability in LLM Agent Teams”, arXiv v1, September 4, 2026.
- Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang, “LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models”, Findings of NAACL, 2025.
- Po-Ya Angela Wang et al., “Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions”, 2026.
- Ray Reagans, Linda Argote, and Daria Brooks, “Individual Experience and Experience Working Together”, Management Science, 2005.
- Robert S. Huckman, Bradley R. Staats, and David M. Upton, “Team Familiarity, Role Experience, and Performance”, Management Science, 2009.
- Jia Li and Daan van Knippenberg, “The Team Causes and Consequences of Team Membership Change: A Temporal Perspective”, Academy of Management Annals, 2021.
- Yufeng Du et al., “Context Length Alone Hurts LLM Performance Despite Perfect Retrieval”, Findings of EMNLP, 2025.