Assume Every Call Happens Twice: Idempotency Is the API-First Discipline AI Breaks First
Photo by Alina Grubnyak on Unsplash
Most API-first programs were designed around a polite assumption: the caller is a piece of software written by someone who read the documentation, handles a 500 sensibly, and calls create order exactly once. Agentic consumers break that assumption on day one. A model-driven caller times out, loses the response, re-plans, and calls again — sometimes with slightly different wording, sometimes from a different sub-agent in the same fan-out. The result is not a model quality problem. It is an interface design problem that your estate has been quietly deferring for a decade.
The retry loop is now a first-class API consumer
The volume shift is already visible. Postman's State of the API 2025 reporting describes API strategy converging with AI strategy, and Kong's GenAI enterprise research found a striking gap: the overwhelming majority of developers now use generative AI, while only around a quarter say they design APIs with AI agents in mind. Gartner, in its 2026 predictions, goes further and expects the bulk of B2B buying to be agent-intermediated within a few years. Whatever you think of the timeline, the direction is clear: a growing share of traffic hitting your order, pricing, and entitlement endpoints will come from callers with non-deterministic retry behaviour.
That makes the retry loop a consumer in its own right — one with no support ticket, no release calendar, and no patience. If your contract does not define what happens on the second identical call, the agent framework defines it for you.
Idempotency keys are a contract term, not an implementation detail
Payments figured this out first, which is why the Agentic Commerce Protocol maintained by OpenAI and Stripe carries idempotency semantics in the spec rather than in a vendor FAQ. Enterprises should copy the posture, not just the header. Idempotency is a promise about business effects: the same logical intent produces one durable outcome, no matter how many times it crosses the wire.
- Key on intent, not on payload. Hashing the request body fails the moment an agent rephrases a free-text field. The key must identify the business intent and travel with it across hops.
- Publish the window. How long do you honour a key — minutes, days? Undocumented retention is a duplicate generator in a slow-moving incident.
- Return the original result, not a conflict. A 409 sends a planner into improvisation. Replaying the first response ends the loop.
- Classify every operation. Safe, idempotent, or effectful. Effectful operations without keys should not be exposed to autonomous callers at all.
- Keep an effect ledger. Record what was actually done, keyed by intent, so reconciliation is a query rather than an archaeology project.
Events make the duplicate visible; APIs only make it expensive
This is where event-driven design earns its place beyond the usual decoupling argument. A synchronous API can suppress a duplicate call, but it cannot tell the rest of the estate that a duplicate was attempted. An event stream can. Carry the intent key into the published event, deduplicate at the consumer, and every downstream system — finance, fulfilment, CRM — inherits the same guarantee instead of reimplementing it badly five times. The stream becomes the shared memory of what actually happened, which is also what gives you a defensible audit trail when an agent's decision is challenged.
Where this gets hard: packaged cores
In SAP and other packaged landscapes, the effectful operations usually sit behind interfaces that predate the idea of an unreliable caller. The pragmatic pattern is an intent layer in front of the core: accept the keyed intent, deduplicate, translate to the standard API, and publish the outcome as an event. It keeps the core clean while giving agents the retry semantics they need. We work through this trade-off regularly in SAP integration engagements, and the same approach underpins how we design API and event platforms for mixed estates.
What to do next quarter
Do not start with a platform decision. Start with an inventory: list every effectful endpoint an agent could plausibly reach, and mark the ones that would double-charge, double-ship, or double-email if called twice. That list is your AI-readiness backlog, and it is usually shorter and more tractable than the architecture diagram suggests.
The organisations that will absorb agentic traffic gracefully are not the ones with the most endpoints or the newest broker. They are the ones whose interfaces assume the caller is unreliable, because from now on, it generally is.
Frequently Asked Questions
Isn't idempotency just a backend concern that agents shouldn't need to know about?
No — the key has to originate with the caller or an intent layer close to it, because only the caller knows that two requests represent the same business intent. If the server tries to infer duplication from payload similarity, it will fail whenever an agent rephrases or re-plans. Treat it as a published contract term with a documented key format and retention window.
How does this differ from exactly-once delivery in messaging?
Exactly-once delivery is a transport-level guarantee that is notoriously hard to achieve end to end. Idempotency sidesteps it by making duplicate delivery harmless at the business-effect level: at-least-once delivery plus idempotent handlers is both simpler and more robust than chasing exactly-once semantics in the broker.
Where should the intent layer live in a packaged SAP landscape?
Usually outside the core, in the integration or API management tier, so the standard interfaces stay untouched and upgrade-safe. It accepts keyed intents, deduplicates against an effect ledger, calls the standard SAP API, and publishes the outcome as an event. That keeps clean-core commitments intact while giving autonomous callers the retry semantics they require.

Comments (0)
No comments yet. Be the first to comment!