When an Allowed Purchase Rests on a Bad Claim
Cryptographic mandates can prove that an agent purchase matches the transaction fields and constraints they encode, but they cannot by themselves establish that the model chose that action from trustworthy evidence; structural violations and semantically plausible choices require different controls.
In one of the synthetic shopping cases in a new security study, the dangerous part is not a forged payment or an altered cart. The agent sees two products it is allowed to buy. Merchant-controlled text supplies a factual-sounding reason to prefer the more expensive one. The agent accepts the premise and changes its recommendation.
The resulting purchase can still look impeccably orderly. The item exists. The cart matches the item the agent selected. The payment details can be internally consistent. The cryptographic objects can be valid. Nothing in the final transaction necessarily says, this choice was reached because the model trusted a claim it should not have trusted.
That is a different kind of failure from somebody changing an amount after approval. It lives earlier, in the path from a delegated request to the transaction that eventually gets signed.
A September preprint by Yedidel Louck, Amit Dvir, and Ariel Stulman uses Google’s Agent Payments Protocol, or AP2, to make that boundary visible. Their experiments do not show that signatures are useless, or that AP2 lacks intent controls. In fact, the current AP2 specification contains a substantial set of them. The more revealing result is narrower: a system can faithfully verify the relationships it represents while leaving the evidence that caused the agent to choose those relationships outside the verification boundary.
Once an AI agent is allowed to interpret open-ended text between authorization and action, the question is no longer only whether the final object is authentic. It is also which parts of the decision can be made explicit enough to check.
What the mandate actually proves
AP2 is designed to let a user delegate commerce without reducing every purchase to blind trust in a model. Its current version distinguishes two broad modes. When the human is present, the user can directly approve the closed checkout and payment mandates. When the human is absent, the user can approve an open mandate that limits what the agent may later do. The verifier then checks the agent’s closed transaction against those constraints. The protocol’s flow documentation makes that distinction explicit.
Those constraints are not trivial. An open checkout can restrict merchants and acceptable line items. Payment constraints can restrict the payee, payment instrument, amount range, budget, execution date, and related references. The payment mandate is also cryptographically linked to the checkout it pays for. In other words, AP2 already embodies an old security lesson: “the user authenticated” is weaker than “the user authorized this particular action under these particular conditions.”
The specification also draws a line around what it does not define. The checkout and payment mandates are assembled after the shopping agent has determined what task the user wants, and the exact method for determining that task is outside the mandate specification. That is a sensible modular boundary. A payment protocol cannot standardize every conversation, catalog, recommendation, or inference that precedes a purchase.
But the boundary matters because modern agents do not simply fill a form. They read descriptions, tool outputs, merchant messages, shipping details, and other text before they decide what the form should contain.
The September study, Signing the Transaction but Not the Decision, probes that gap with three different attacks. One tries to make the agent retrieve payment credentials belonging to another identity. A second tries to push a product into the cart that does not match what the agent had actually been shown. A third changes the agent’s product choice with a factual-looking claim while keeping the resulting cart structurally consistent with what was displayed.
These are easy to lump together as prompt injection. That description is true but incomplete, because the three failures leave different evidence behind.
Some bad decisions leave a broken edge
Suppose a shopping session belongs to Alice but the agent asks a credential service for Bob’s card. There is a relationship the system can test: session identity should match credential identity. If the agent should never be able to choose the credential owner in the first place, the safer interface is even simpler—bind that lookup to the session and remove the choice.
Or suppose the agent saw listing A but later submits listing B. Again, there is a relationship to test: each cart line should correspond to a listing the agent actually received and was permitted to use.
Louck and colleagues propose a defense they call A-VIP that adds bindings of this kind. In their evaluation, it blocks the first two attack families without reported false positives. The interesting part is not the name of the defense. It is what makes the defense possible: the bad action leaves a structural mismatch that software can express as an invariant.
The third attack is more awkward. The agent sees two legitimate choices. An untrusted statement changes which one it prefers. The final item is still one of the items it saw. The cart is still internally coherent. If the user’s delegation genuinely allowed either product, there may be no wrong SKU, wrong merchant, wrong identity, or wrong payee for a deterministic verifier to reject.
The researchers therefore do something different: they surface the spending change for confirmation rather than claiming that another cryptographic link can decide whether the factual premise was trustworthy.
This is the useful distinction. Some drift is trace-bearing: the agent’s action breaks a relationship the system could have represented and checked. Other drift is trace-free at the transaction layer: the agent makes a plausible choice that satisfies the formal constraints but may still rest on bad evidence or bad judgment.
Cryptography is excellent at the first problem once the relevant relationship has been represented. It cannot, by itself, make an untrusted product claim true.
The headline numbers are not prevalence rates
The paper reports eye-catching success rates in its default setup: 90 percent for the credential attack, 56 percent for the cart attack, and 73.3 percent for the product-selection attack. Those numbers establish that the failure modes can be made high-yield under the tested configuration. They should not be read as estimates of how often ordinary AP2 purchases will go wrong.
The authors released AP2-WhisperBench, which makes that limitation unusually visible. An earlier diversity-holdout version of two attack families reports much lower undefended success rates—32 percent and 14 percent. An intermediate version reports 60 percent and 30 percent. The dataset’s default configuration uses a later set whose wording was optimized for attack yield.
That sensitivity is not a reason to dismiss the result. Security research often asks whether an attacker can reliably create a failure, not how frequently a random input accidentally creates one. But it changes what the evidence supports. The benchmark demonstrates feasible mechanisms under synthetic, selected conditions; it does not measure a field rate for autonomous commerce.
There is another reason to resist a sweeping conclusion. AP2 itself has stronger controls than the caricature “sign whatever the model decided.” Fully populated merchant, line-item, budget, and payee constraints can eliminate classes of bad outcomes before any new mechanism is needed. Some of the study’s first two attacks may therefore tell us as much about least privilege, interface design, and constraint population as about a new cryptographic problem.
That is exactly why the third attack matters. It survives the easy explanation. A user can intentionally leave a choice open because autonomy is the point. Once two choices are both authorized, the remaining problem is no longer simply permission. It is the quality and provenance of the evidence used to choose between them.
This problem did not begin with AI agents
Security engineers have spent decades separating identity from transaction-specific authorization. OAuth’s Rich Authorization Requests, for example, exist because a coarse permission such as “payments” cannot express something like “transfer this amount to this merchant.” The standard lets applications carry structured authorization details such as amount, currency, and creditor.
AP2 extends the same basic instinct into agentic commerce: represent more of the action explicitly, link related objects, and verify those links deterministically.
So the new lesson is not that digital signatures somehow fail to capture meaning. They have never promised to certify every reason behind an action. Nor is prompt injection against AP2 new. A January paper, Whispers of Wealth, had already shown that hostile commerce content could steer an agent.
The more specific change is that an LLM now sits between the user’s delegated goal and the structured transaction. It performs an open-ended semantic transformation while consuming information from sources with very different trust levels. Some of what it decides can be pushed into fields and invariants. Some of it cannot.
An independent August analysis of AP2 v0.2, Beyond the Mandate, reaches a compatible boundary from another direction. It models threats across the protocol lifecycle and finds that signed mandates protect transaction state after signing while pre-authorization interactions and external inputs remain an attack surface. That does not validate every result in the September benchmark, but it makes the location of the problem less dependent on one paper’s attack construction.
The strongest alternative explanation is still important: perhaps the real fix is simply better prompt-injection resistance, source isolation, or safer tools. For the credential attack, that may be most of the story. For a cart that violates an explicit item constraint, existing policy machinery may already be enough. Better models and cleaner interfaces should reduce the attack surface.
But they do not remove the category. As long as an autonomous agent is allowed to make a legitimate choice using open-ended evidence, there will be decisions for which every represented transaction field is valid and the unresolved question is whether the model believed the right thing.
Four questions that look like one approval
An “approved” agent action can collapse several different claims:
- Who acted? Authentication and signatures can establish identity and integrity.
- Was the actor allowed to act? Delegation and capability scope can establish authority.
- Does this action satisfy the constraints we represented? Deterministic checks can enforce merchants, items, amounts, payees, resources, paths, or similar relationships.
- Was this action chosen for reasons we should trust? That may require source controls, provenance, policy, or human judgment.
The fourth question is tempting to solve by keeping more logs. Provenance helps when it exposes a relevant mismatch: which listing the agent saw, which tool result it used, which account a credential came from. But provenance establishes where information came from, not that the information was true. A perfectly signed record of a false merchant claim is still a false merchant claim.
That suggests a more practical design rule than “record everything.” Move security-relevant relationships into deterministic enforcement whenever the system can state them cleanly. Do not let a model choose an identity that the session already determines. Bind authorized items and payees to the objects the user or trusted system actually approved. Mark external content as untrusted and keep it away from privileged instruction channels where possible.
Then be explicit about what remains. If a high-stakes choice cannot be reduced to a reliable structural rule—because both options are allowed and the difference is a judgment about evidence—either give the agent a trustworthy source for that judgment, apply a policy threshold, or bring the human back into the loop.
There is a cost to all three. More constraints reduce flexibility. More provenance consumes storage and reviewer attention and can anchor reviewers on the agent’s own story. More confirmations make autonomy less autonomous and can deteriorate into reflexive clicking.
That tradeoff is why the synthetic product-selection case is more useful than a story about a forged signature would have been. The interesting transaction is the one in which the cryptography works. The item is valid. The cart is coherent. The user’s delegation may even permit the purchase. What remains is a question the signature was never designed to answer: why did the agent choose this one?
For a consequential purchase, the system has to know whether that question can be converted into a check. If it cannot, the honest control is not another seal on the same cart. It is a decision about which evidence the agent may trust—or whether this particular choice should return to a person before money moves.
Sources
- Yedidel Louck, Amit Dvir, and Ariel Stulman, “Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2”, September 10, 2026.
- AP2 v0.2 specification, including the protocol’s mandate model and verification boundary.
- AP2 flows, including Human Present and Human Not Present modes.
- AP2 Checkout Mandate and Payment Mandate constraint documentation.
- AP2-WhisperBench, released benchmark and dataset card.
- Roei Aviv et al., “Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)”, August 24, 2026.
- Tanusree Debi and Wentian Zhu, “Whispers of Wealth: Red-Teaming Google’s Agent Payments Protocol via Prompt Injection”, January 30, 2026.
- RFC 9396, OAuth 2.0 Rich Authorization Requests, May 2023.