Fresh Facts, Stale Plans
When model-generated plans persist across state changes, their validity can expire independently of the freshness of the executor's current context; safe reuse therefore requires some representation of the state that licensed the plan, balanced against the cost and incompleteness of dependency tracking.
Terraform has an error message for a peculiar kind of obsolescence.
A saved plan is a durable object. It records the work Terraform intends to perform against a particular infrastructure state, so someone can inspect the plan and apply it later. But Terraform Core also stores metadata about the state from which that plan was made. If another operation changes the state’s serial number before the plan is applied, the local backend can refuse it with a blunt diagnosis: “Saved plan is stale.”
Nothing has vanished. The plan is still there. The current state is still there. The problem is that they no longer describe the same moment.
That distinction is becoming relevant to AI agents for a reason that has little to do with memory in the usual sense. Agent systems increasingly preserve plans across model calls, hand them to other agents, resume them after tools run, or store them as durable work artifacts. Once a plan survives the instant in which it was generated, the world can change underneath it.
A new controlled study demonstrates the resulting failure in unusually clean form. Its lesson is narrower than “agents need fresher context.” The executor can already have the new context.
The old plan can still be the thing telling it what to do.
The executor knew the new requirement
In Fresh Memory, Stale Plans, Evan Chen, Shiqiang Wang, and Christopher G. Brinton construct distributed-agent workflows in which one agent plans from a public requirement and another revises that requirement before a consequential action. The revised record reaches the executor before the action occurs.
Suppose the planner reads requirement r3 and produces plan p3. Then r4 becomes the authorized requirement. Updating the executor’s shared state to r4 changes what the executor can see. It does not, by itself, change what produced p3.
The authors stage that race in 30 live workflows spanning reservation, fulfillment, and deployment tasks. Each task deliberately receives one post-plan revision in the dangerous interval before the protected action. A freshness-only policy still issues the obsolete primary action in all 30 tasks. Exact-lineage policies, including the authors’ PlanFence protocol, detect the mismatch and prevent that defined invalid action.
They also replay the protocol with model decisions fixed, removing variation in generation from the comparison. Under those staged schedules, freshness-only policies issue 330 obsolete actions in 330 opportunities; exact-lineage policies issue none.
Those are protocol results, not prevalence estimates. The experiment is designed to create the race. It does not show that production agents ignore changed requirements 100 percent of the time, or even that stale-plan execution is common. The paper is an arXiv v1 submitted on September 3, 2026, and has no independent replication yet. Its own exploratory audit found the defining pattern in 15 of 30 workflows, but the authors explicitly decline to treat that as a natural incidence estimate.
What survives those qualifications is the conceptual separation the experiment makes visible.
There are two different questions hidden inside the phrase “up to date.”
One is about state: what is true now?
The other is about derivation: what was true when this reusable plan was made?
A system can answer the first correctly and still have no answer to the second.
An old problem acquires a new object
Computer systems have been dealing with expiring derivations for decades.
Optimistic concurrency control lets a transaction proceed and then checks whether the data it relied on changed before the transaction commits. Build systems reuse an old output only while the inputs that matter to it remain valid. GitHub can dismiss an approval after the approved diff changes because the approval belongs to a particular state of the change, not to the pull request for all time. Classical AI planning has likewise treated plan validity as conditional on a world that can move.
So dependency invalidation is not the new idea.
The new object is the model-generated plan.
Until recently, much LLM reasoning was ephemeral. A prompt arrived, a response was produced, and the reasoning disappeared into the interaction. Agent systems increasingly make intermediate reasoning durable. A plan may be stored in shared memory, represented as a task, handed across an agent boundary, resumed tomorrow, or used to authorize a tool call after other work has changed the repository.
Persistence is useful precisely because the system does not have to derive everything again. But persistence also creates a new lifecycle question: when is reuse no longer justified?
That question cannot be answered by copying the newest facts into memory. A new fact can coexist with an old conclusion.
The easy check hides the hard problem
PlanFence approaches this as a lineage problem. The plan records parent versions for public state, and each protected action has a declared set of public inputs that can affect it. Immediately before acting, the executor compares the versions associated with the plan against the authoritative current versions. A mismatch triggers replanning; incomplete evidence or another change can block the action.
The attraction is obvious. If a deployment action depends on three records, it should not need to invalidate itself because an unrelated fourth record changed. In the paper’s controlled systems experiments, scoped validation becomes more attractive as churn and irrelevant shared state grow. Proactive synchronization performs better in most of the measured low-churn settings, and the advantage of scoping shrinks when actions depend on much of the shared keyspace.
That tradeoff exposes the real engineering problem.
Comparing version IDs is easy. Knowing which versions mattered is not.
PlanFence assumes that trusted application code provides a complete dependency set for the action. Miss a relevant input and the protocol cannot validate it. Declare nearly everything relevant and the system recovers safety by checking nearly everything, giving up much of the value of fine-grained invalidation.
Software build systems know this failure well. Missing dependencies can leave an apparently current output built from an obsolete input. Redundant dependencies can make a build correct but unnecessarily expensive. The difference is that conventional software often leaves machine-readable traces—imports, file reads, schemas, compiler dependencies. A language model may absorb a premise from a paragraph, an example, a convention, or the relationship between two files without recording a neat edge that says, “this fact licensed that decision.”
A plan can therefore have excellent provenance for the facts the system remembered to name and still be falsely safe about the one it forgot.
Terraform is safe by being crude
Terraform is useful here not because it has solved semantic invalidation, but because it shows the opposite strategy.
For local saved plans, Terraform Core compares the state lineage and serial embedded in the plan with current state metadata. A serial mismatch is enough to reject the plan. That is conservative. It does not ask whether the particular change actually affects the actions in the plan.
A long-running Terraform issue illustrates the cost. Users documented cases in which a changing data source advanced the state serial even when there were no actionable resource changes. A previously saved plan could then be rejected as stale anyway.
Terraform was not wrong about the version. It was deliberately agnostic about the meaning of the version change.
That is the tradeoff agent systems inherit. A whole-repository rule is simple: if the commit moved, replan. It is also likely to throw away valid work on a busy branch. File-level invalidation is cheaper but can miss cross-file relationships. Region-level invalidation is more precise but must survive edits and refactors. Semantic invalidation promises fewer pointless replans, but only if the system can reliably identify the premises that mattered to the model.
There is no free granularity.
The strongest rival to fine-grained lineage is therefore not “do nothing.” It is invalidate more broadly. In many systems, global replanning or conservative version binding may be simpler, easier to audit, and safer than maintaining a sparse dependency graph whose omissions are difficult to detect. In other systems, especially high-churn ones with expensive plans and narrow action dependencies, that conservatism may be intolerably wasteful.
The right architecture depends on the cost of being wrong in both directions: silently reusing invalid reasoning versus recomputing reasoning that was still valid.
Persistence changes what a plan is
This reframes a familiar discussion about agent memory.
Most memory systems are evaluated by what they retain, retrieve, summarize, or supersede. Persistent reasoning introduces another kind of object. A plan, approval, analysis, or test interpretation is not merely a fact to remember. It is derived work—an artifact produced from earlier facts and later reused because the system assumes the derivation still holds.
For derived work, storage is only half the lifecycle.
The other half is invalidation.
That does not mean every private thought needs to be versioned. PlanFence itself excludes private prompts and hidden reasoning from its public lineage guarantee. Nor does its validation form a transaction across multiple owners and the eventual external side effect. The paper assumes benign authoritative owners, authenticated transport, immutable parent links, and complete declared action dependencies. Those are meaningful boundaries, especially for software work where relevant premises can be implicit.
It also remains possible that capable coding agents often notice changed premises and revise their plans without any explicit lineage mechanism. The controlled study shows that fresh state does not guarantee fresh derivation. It does not establish how often explicit invalidation will outperform spontaneous replanning in real repositories.
But once a workflow deliberately preserves a model-generated plan beyond the state that produced it, one assumption should disappear: that refreshing context refreshes the plan.
Terraform’s stale-plan error is crude, but it is precise about the relationship it refuses to assume. The saved plan still exists. The newer state still exists. Terraform simply will not pretend that one still authorizes action against the other.
Persistent agent systems will need some version of the same answer—whether coarse or fine-grained—before an old plan becomes a new action.
Sources
- Evan Chen, Shiqiang Wang, and Christopher G. Brinton, “Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory”, arXiv v1, submitted September 3, 2026.
- HashiCorp Terraform Core, current local-backend saved-plan validation.
- HashiCorp, “Run modes and options in HCP Terraform”, saved-plan behavior.
- HashiCorp Terraform issue #27827, “Subsequent applications with no changes cause state’s serial to increment, causing stale plans”, 2021.
- GitHub Docs, “Available rules for rulesets”, stale approval behavior.
- H. T. Kung and John T. Robinson, “On Optimistic Methods for Concurrency Control”, ACM Transactions on Database Systems, 1981.
- Maayan Shvo, Toryn Q. Klassen, and Sheila A. McIlraith, “Resolving Misconceptions about the Plans of Agents via Theory of Mind”, ICAPS 2022.