Skip to content

What Must Travel With an Agent’s Memory

Related:Cards ↗

For long-lived agents, preserving content is not enough: provenance, unresolved safety evidence, and execution authority are separate state that may need to survive different boundaries for different lengths of time.

In August, researchers found a command on Clerk’s legitimate website that told a coding agent to run npx clerk-next-fix-auth-protection.

The name looked plausible. The problem was that the npm slot it named had once been empty. By the time the researchers checked, someone had claimed it and was serving malware. Clerk later fixed the documentation; whether the malicious package caused real infections remains unknown. Ars Technica reported the case on August 27.

Nothing about that episode requires an agent to have a long-term memory. It is a supply-chain and provenance failure. But it exposes a useful fact about durable information: the text can survive while the conditions that made the text safe change underneath it.

That becomes a harder problem when an agent’s own state survives for hours, days, or across restarts. A plan can be summarized. A tool result can be copied into a note. A warning can be compressed away. A recommendation can remain while the evidence about where it came from disappears.

The ordinary memory question is: What information will help later?

A long-lived agent also needs a different question: What must remain true about that information while it lasts?

A warning can disappear too soon

One reason agent memory exists at all is that useful work outlives a context window. The system has to decide what to retain, summarize, retrieve, or forget. For ordinary task context, forgetting can be helpful. SF-AMS, for example, reports gains on its evaluated long-context benchmarks by keeping high-utility memories and filtering redundant or low-value material.

Security can reverse that logic.

An August 27 preprint called Safety Does Not Compose constructs an attack whose decisive evidence is deliberately split across iterations. Each local fragment is made to look like benign behavior. A safety monitor that sees only one local trajectory therefore cannot distinguish the attack from the benign case better than its own false-positive rate. The missing evidence is not hidden from a weak model. It is absent from the model’s input.

A monitor with persistent cross-iteration state can distinguish the joint pattern.

The result is strong but narrow. The paper does not show that most deployed attacks have this structure. It establishes a boundary: if the evidence required for a safety judgment lasts longer than the state available to the judge, making the judge smarter cannot recover evidence it never receives.

The paper also tests an intuitive compromise: let old risk gradually fade. Under its geometric-decay model, that creates a finite cooling-off period. A patient adversary can wait until the earlier signal matters little enough and then continue. The authors do not conclude that every suspicious signal should live forever. Their own design selectively retains grounded attack evidence and a loop-level flag while allowing other signals to reset.

A separate May preprint, MAGE, provides a more empirical version of the same idea. It maintains a dedicated safety-focused “shadow memory” across an agent’s execution trajectory and checks pending actions against it. In the two long-horizon threat families the authors evaluate, the added safety state substantially reduces attack success while preserving meaningful benign utility. Those experiments do not prove the newer paper’s theorem, and they do not establish the same result across arbitrary restarts or multi-agent systems. They do show why a security record can deserve a longer life than ordinary task detail.

The lifetime of a warning should be set by the hazard it helps identify, not merely by how useful the warning seems to the current task.

The same durability can preserve an attack

If old information can protect an agent because it influences future decisions, old information can also compromise one for the same reason.

AgentPoison, published at NeurIPS 2024, poisoned agents’ long-term memories or knowledge bases. Across the agent settings it tested, the paper reported average attack success above 80 percent with a poison rate below 0.1 percent and less than one percent change in benign performance.

A June 2026 preprint, From Untrusted Input to Trusted Memory, studies the problem more systematically. It identifies four routes by which adversarial material can enter agent memory and nine structural vulnerabilities across the systems it examines. Its benchmark results indicate that agents that write and retrieve memory more aggressively can also expose a larger attack surface.

These results complicate any simple prescription to preserve more history. Persistence is neither the defense nor the vulnerability. It extends the life of whatever influence the system has chosen to carry forward.

That produces the central tension. One old fact may need to remain consequential because a future action completes a dangerous pattern. Another old fact may remain useful as a suggestion while never deserving more authority than it had when first observed.

Those are not two settings on one memory knob.

A suggestion and a permission are different state

This distinction sounds new only if “memory” is treated as a special AI category. In systems security, separating data from authority is old work.

CaMeL, introduced in 2025, applies that lineage directly to tool-using language-model agents. Instead of asking the model to perfectly recognize malicious instructions, it separates trusted control flow from untrusted data and uses capability-style checks to constrain what information can cause a tool action.

That is the strongest rival to the idea that agent memory needs a new theory. Provenance, information-flow control, capabilities, and security logs did not appear because language models acquired context windows. Much of the problem is classic security engineering in a system whose intermediate state is often natural language.

The agent-specific question is narrower: do those control properties survive the same transformations and boundaries that the agent’s work survives?

Two fresh preprints make that temporal problem explicit. When Tool Outputs Become Commands separates action induction from execution authorization. An untrusted webpage can legitimately suggest a useful command; that does not mean the webpage should acquire permission to make the command run. Its SARA architecture therefore records the origin of action-inducing information and checks execution separately against the user’s objective and already authorized work. Its “No-History-Promotion” rule says that seeing the same origin again in history cannot, by itself, upgrade that origin into execution authority.

In four primary AgentDojo and AgentDyn settings, the authors report attack success no higher than 0.63 percent. That is promising systems evidence, not proof that repeated exposure ordinarily makes deployed agents more trusting. The rule prevents an authority-laundering pathway by construction; it does not establish the prevalence of that pathway in the wild.

Another August 27 preprint, SPA, stores persistent execution results as labeled artifacts and exposes semantic metadata, rather than the raw old payload, to later planning. Integrity labels survive across queries. The authors also report a security-utility tradeoff from stricter enforcement. SPA is currently an arXiv preprint submitted to USENIX Security 2027, not an accepted conference paper, and its AgentDojo-family evidence is not an independent replication of SARA.

The useful conclusion is smaller than “typed memory solves agent security.” A long-lived system has several kinds of state that happen to travel together unless the architecture deliberately keeps them apart.

The sentence is one kind. Its origin is another. An unresolved warning is another. Permission to cause a side effect is another still.

Different clocks

Once those properties are separated, the memory policy stops looking like a single retention rule.

A debugging detail may lose value quickly. A provenance record may need to remain attached as long as the derived note can influence protected work. A warning may need to survive until the condition it describes is explicitly cleared. A permission may expire even while the instruction it once authorized remains accurate.

This also explains why “never forget” is the wrong security lesson.

Permanent suspicion can overblock legitimate work. Strict integrity rules can reduce utility. More retained state creates storage, privacy, and review costs. Research on strategic forgetting is useful precisely because indiscriminate persistence can make an agent worse at its task.

The better rule is conditional: retain a security fact for the period in which the hazard can depend on it, or until a justified event clears it. Do not treat survival itself as new evidence of trust.

That implies at least two different kinds of forgetting. A system may discard stale task detail while retaining unresolved safety state. It may also keep a useful instruction while preserving the fact that the instruction came from a source with no authority to trigger execution.

The hard part is not storage. It is transformation.

What gets lost in a summary

A long-running agent rarely preserves its past as an untouched transcript. It summarizes, compacts, indexes, merges, retrieves, and hands work to another process.

Imagine an untrusted page recommends a command. The raw page is marked as untrusted. Hours later, the system compresses the session into one sentence: “To repair the deployment, run command X.” If the summary carries the recommendation but drops its origin, the system has changed the security meaning of the information without changing the instruction itself.

The inverse failure is just as important. Three earlier actions may look harmless separately and become dangerous only in combination. If the work summary retains the current objective but discards the earlier actions, the next safety check cannot see the pattern.

That is what persistent agent systems should test explicitly. Can a summary be traced to the observations that produced it? Does a restart preserve unresolved safety state? Is there a documented event that clears a warning? Does a tool action require authority independent of the text that recommended it? If a bad memory write is discovered, can the affected state be identified and rolled back?

Those questions are less glamorous than “How much memory should an agent have?” They are also closer to the engineering problem.

The command can still look right

Return to the Clerk command.

The line appeared on a legitimate site. The package name looked plausible. What changed was the ownership and behavior of the thing the name resolved to.

For a persistent agent, an analogous separation can happen internally. The note survives; the source marker does not. The instruction survives; the permission expires. The task summary survives; the warning that made the next step unsafe is gone.

A durable record is useful only if the properties needed to interpret it survive the journey too.

The question is not simply what the agent remembers. It is what the system refuses to forget about what it remembers.

Sources