What a Successful Tool Call Can’t Tell You
For tool-using agents, call-local success is only evidence about one invocation; workflow correctness also depends on effect identity, authoritative outcome, reversibility, dependencies, ordering, and visibility that may not be standardized at a shared tool boundary.
Amazon’s engineering guidance poses a small distributed-systems problem that becomes unpleasant almost immediately.
Imagine asking EC2 to launch exactly one server. The request leaves your machine. Then the connection times out. No response comes back.
Did the server start?
If it did not, retrying is sensible. If it did, retrying may create the second server your “singleton” workload was designed to avoid. The request itself was valid in both cases. The uncertainty sits somewhere else: between what happened in the world and what the caller managed to learn about it.
AWS’s answer is not to make the caller guess better. Its idempotent API design lets a client attach a unique token to the intended operation. A retry carries the same token, so the service can recognize it as another attempt at the same action rather than as a new, coincidentally identical action. AWS also points out a less glamorous requirement underneath the trick: recording the token and performing the mutation have to be coupled atomically, or the bookkeeping can disagree with the resource.
This is old distributed-systems engineering. It becomes newly important when an AI agent is allowed to string together tools that were built by different people, against different services, with different ideas about what a successful call proves.
A model can choose the right tool, supply valid arguments, and receive a clean result. The workflow can still end in the wrong external state.
A result is not the state of the world
A paper submitted on September 14 by Artem Trofimov and Boris Novikov gives this problem an agent-specific vocabulary. In When Tool Calls Succeed but Workflows Fail, they separate two things that software often lets us blur together: an effect, meaning a change that occurred in an external system, and an observation, meaning what the runtime learned about that change.
The distinction is easy to miss on a happy path. A tool runs, the service changes state, a response arrives, and the agent continues. Event and observation appear to travel as one package.
Failures pull them apart.
A payment can go through while its acknowledgment disappears. A speculative branch can send an email before another branch wins. A workflow can cancel a reservation but leave behind a message that caused someone else to act. Two agents can publish incompatible updates to the same external resource without either seeing the other’s move first.
Trofimov and Novikov organize these cases into eight anomaly patterns. The list is useful less as a new theory of transactions than as a map of where an agent runtime runs out of facts. Some failures require knowing whether one logical operation already took effect. Others require knowing whether an effect can be staged or compensated. Still others require stable resource identity, dependency information, an ordering rule, or control over when outsiders can see provisional state.
The causal variable is not how intelligent the model is. It is what the boundary lets the runtime establish.
That changes the meaning of a tool log. A line that says success may tell you that one invocation completed according to one service’s contract. It does not, by itself, tell you that the resulting effect belongs in the workflow that eventually survives.
Consider cancellation. If an agent books a non-refundable ticket and later abandons the trip, the booking service may have behaved perfectly. So may the agent, in the narrow sense that it called the intended operation with legal parameters. The failure appears one level up: an effect from a workflow that no longer exists has escaped into the world.
Compensation helps, but it does not erase history. An email can be withdrawn after a recipient has acted on it. A reservation can be canceled after another system has scheduled work around it. Once an outside observer has reacted, undoing the original action does not automatically undo the reaction.
That is where a familiar database analogy stops being merely decorative. Transactions work because the participating system controls enough of the state and protocol to decide what becomes visible, what commits together, and what can be rolled back. An agent calling third-party tools may control none of those things.
The same request is not necessarily the same operation
AWS’s client token reveals one of the missing pieces with unusual clarity.
Two requests with identical parameters are not always duplicates. If I ask for one EC2 instance twice, I may want two instances. If a network timeout causes an SDK to retransmit the same logical request, I may want only one. The bytes can look almost identical while the intent differs.
The service needs an identity for the intended operation that survives retries.
Stripe exposes the same idea in its idempotency-key contract. A client can attach a key to a create or update request; for API v1, Stripe saves the status code and response body produced by the first execution and returns that result to later requests carrying the same key, including when the first response was a server error. The important fact is not that a retry resembles the first request. It is that both attempts name the same logical action.
This solves a narrow but important class of uncertainty. It does not solve an entire workflow. A deduplicated payment can still be inconsistent with a branch that later loses. A perfectly queryable booking can still become premature if it is made visible before approval. Two individually idempotent operations can still conflict on shared state.
The lesson is therefore not “add idempotency.” It is that recovery mechanisms consume semantics. A runtime can only make a safe decision from facts the lower layer gives it or that an application-specific adapter already knows.
What a shared tool protocol actually says
This is the interesting part of the recent paper because the underlying systems ideas are old.
Trofimov and Novikov examined the standard annotation vocabulary exposed by reachable remote servers in the official Model Context Protocol registry. Their July 27 snapshot yielded 98,291 tool definitions from 4,838 remote targets that returned at least one tool. They did not call those tools. They inspected what the interface declared.
Seventy-four percent serialized at least one of MCP’s four familiar behavior hints; 61.7 percent serialized all four. Those fields can tell a client that a tool is intended to be read-only, destructive, idempotent, or open-world. But even a fully populated set does not provide an idempotency key, an authoritative outcome query, a compensation contract, a prepare/commit mechanism, a commutativity rule, or a way to say which effects depend on which provisional state.
That does not mean 98,291 tools are unsafe. The study never measured their real failure behavior, and a tool may be backed by a service with much stronger guarantees than its MCP metadata advertises. The current MCP specification also explicitly treats tool annotations as optional behavior hints that clients should consider untrusted unless the server itself is trusted. The protocol is not promising a transaction layer and then failing to deliver one.
Nor is MCP unable to carry state. Its current Tasks extension gives a participating server a durable task handle that a client can poll, and advises clients to persist task IDs so they can resume after a crash or disconnect. Current tool-design guidance also recommends explicit handles for state that spans calls rather than hidden connection-local state.
Those are meaningful counterexamples to a broad claim that a shared tool protocol can only return success or error.
They also sharpen the narrower gap. A durable task status tells the client what the server says about a task. It does not automatically prove that a separately managed payment, booking, deployment, message, or other business effect is atomically coupled to that task record. A state handle can name a basket without specifying whether two mutations commute or whether an abandoned branch’s side effects can be neutralized.
The question is not whether a protocol can transport more JSON. It is which semantics are standardized well enough that a runtime built by someone else can safely rely on them.
Agent runtimes are rebuilding the missing facts locally
Recent agent research already makes this boundary visible from the other direction.
Systems such as Atomix and Cordon add transaction-like machinery around tool use: staging some outward effects, tracking enough information to compensate others, or delaying release until a larger unit of work has been validated. Their existence is evidence that the problem matters. It is also evidence against claiming that agents have just discovered transactions.
These systems work by knowing more about their tools than a bare call result says. Atomix, for example, categorizes effects and relies on declared properties to decide what can be buffered, compensated, or released. Cordon likewise stages outward-facing actions inside a task-level transaction model. The details differ, but the architectural shape is the same: stronger guarantees come from stronger contracts about the operations below.
That is why a wrapper cannot always rescue an arbitrary black box. Suppose a request times out after a side effect. If the tool offers no stable operation identity, no outcome query, and no safe reissue semantics, the wrapper may face two possible worlds that look identical from above: the effect happened, or it did not. More reasoning cannot recover information that never crossed the boundary.
The wrapper can wait. It can ask a human. It can accept a weaker guarantee. What it cannot do is infer an authoritative fact from an interface that exposes none.
Stronger guarantees have a price
It would be easy to end here with a universal recommendation: stage everything, serialize every conflict, and make every effect transactional.
That would be another category mistake.
A 2026 CIDR paper on correctness in data-oriented workflow systems shows the tradeoff in a conventional, non-agent setting. In its e-commerce experiments, transactional workflows performed better under low contention. Under hot contention or long-running steps, saga-style workflows could deliver better throughput and avoid the aborts caused by holding transactional locks across more work.
The same tension appears inside the agent paper. Some interactions cannot be fully staged because the outside party’s response is the information the workflow needs next. An agent negotiating an offer cannot learn how the counterparty reacts without exposing the offer. In that setting, “make nothing visible until the workflow is done” is not a safety policy. It is a way to make the workflow impossible.
So there is no single strongest setting to apply everywhere. The useful distinction is between unknown semantics and chosen semantics.
A team may deliberately accept compensation instead of atomicity, or early visibility instead of staging, because latency, interactivity, or throughput matters more. That is a design decision. It is different from discovering during a failure that the runtime never knew which guarantees its tool supplied.
The audit trail has to cross the boundary too
This changes what it means to inspect an agent run after the fact.
A transcript that records the model’s prompt, its tool name, its arguments, and the returned result can be complete at the conversational level and still be incomplete at the systems level. To reconstruct an external effect, a reviewer may also need the logical operation identity, the authoritative external outcome, the resource that changed, the workflow branch that issued it, any dependency on provisional state, and what happened during retry or compensation.
Not every tool needs every field. A read-only weather lookup has different stakes from a payment or deployment. A simple synchronous mutation may need only stable retry identity. A speculative multi-tool workflow may need staging and dependency information. An open-world action may have consequences no runtime can fully retract.
The point is not to turn every tool schema into a database textbook. It is to stop treating the call boundary as if it already contained facts it does not contain.
AWS’s ClientToken is a small example of what changes when a system makes one of those facts explicit. The token survives a timeout. It lets the service recognize that a second network request is another attempt at the first intention, not a new intention that happens to look the same.
For an agent runtime, that distinction can be the difference between retrying safely and making the world twice.
Sources
- Artem Trofimov and Boris Novikov, When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent–Tool Boundary
- Malcolm Featonby, AWS Builders’ Library, Making retries safe with idempotent APIs
- Stripe API Reference, Idempotent requests
- Model Context Protocol, Tools (current draft)
- Model Context Protocol Tasks Extension
- Michael Stonebraker, Xinjing Zhou, Peter Kraft, and Qian Li, Consistency and Correctness in Data-Oriented Workflow Systems
- Bardia Mohammadi et al., Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows
- Zheng Chen et al., Cordon: Semantic Transactions for Tool-Using LLM Agents