← Back to Blogs
September 28, 2026By Cosmoneural Insights

Replay Is the Feature: Why AI-Ready Architectures Are Judged by Retention, Not by Broker Choice

Event-Driven ArchitectureAPI StrategyEnterprise AI
Replay Is the Feature: Why AI-Ready Architectures Are Judged by Retention, Not by Broker Choice

Photo by Joshua Sortino on Unsplash

Ask an enterprise architecture team whether they are event-driven and you will get a platform answer: Kafka, EventBridge, Solace, maybe an integration suite with a topic model. Ask whether they can replay every order, price change and status transition from last Tuesday, in order, into a new consumer without touching the source systems, and the room goes quiet. That gap matters more now than it did three years ago. AI systems — retrieval pipelines, forecasting models, agents that need to understand how a case got into its current state — are history consumers. They do not just want the current value of a field. They want the sequence that produced it. An architecture that streams but does not retain is an architecture that can serve dashboards and will keep failing AI workloads.

APIs answer "what is true now." Events should answer "how did it get that way."

API-first and event-driven are not competing philosophies; they are two halves of one contract surface with different jobs. A well-designed synchronous API gives an agent or service authoritative current state and a transactional write path with clear idempotency semantics. Events give the temporal dimension: the ordered, timestamped record of change that lets you reconstruct state at any point in the past. The failure mode we see repeatedly is teams buying an event platform and then using it as a faster API — fire-and-forget notifications with a seven-day retention window and no schema history. You end up paying for infrastructure that encodes none of the thing you actually needed.

The practical test is simple: pick a model or agent in production and ask what its training or grounding data would look like if you had to rebuild it from scratch next quarter. If the answer involves nightly extracts, database snapshots and a data engineer reverse-engineering business logic, the event layer is not doing its job.

Retention policy is an AI decision, not an infrastructure line item

Retention usually gets set by whoever was worried about storage cost during the platform rollout. Then a lakehouse or an ML pipeline starts replaying history rather than consuming only the live tail, and the seven-day default becomes an architectural constraint nobody agreed to. Tiered storage — where brokers keep hot segments locally and push older segments to object storage, as introduced in Apache Kafka through KIP-405 — has made long retention economically reasonable, but most organisations have not revisited the policy since. The design questions worth reopening:

  • Replay horizon per domain, not per cluster. Order and payment events may need years for reconstruction and audit; telemetry may need days. A single global retention setting guarantees you are simultaneously overpaying and under-retaining.
  • Compacted state topics alongside full history. Log compaction gives fast rehydration of current state; the raw log gives the causal sequence. AI workloads usually need both, for different stages of the pipeline.
  • Schema versioning that survives replay. Replaying two-year-old events through today's consumer is only safe if schemas are versioned and evolution rules are enforced. AsyncAPI definitions and an event catalogue make this governable rather than tribal.
  • Point-in-time permission context. Replaying history through a retrieval pipeline can leak data that a user was never entitled to see. Entitlement at time-of-event needs to be part of the record, not resolved at read time.
  • Deletion and consent obligations. Long retention collides with erasure rights. Decide early whether you tokenise, crypto-shred, or keep personal data out of the durable log entirely.

Design the write path for non-human callers

Gartner has projected that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from under 5% in 2025. Whatever the exact number, the direction is clear: a growing share of API traffic will come from callers that do not read documentation, retry aggressively, and occasionally misinterpret a field name. That changes what "good API design" means. Idempotency keys stop being optional. Error responses need to be machine-actionable, not prose. Rate limits and quotas need to be scoped per agent identity, not per application. And every write that an agent can perform should emit an event carrying the actor, the intent and the reversibility characteristics of the action — because that emitted trail is what makes autonomous behaviour auditable after the fact.

What this looks like in an SAP-heavy landscape

Most enterprises are not building this on a blank slate. They are building it around ERP and commerce platforms where the transactional truth lives and where custom extension has historically been the path of least resistance. The pattern that works: treat the core as the authoritative write path exposed through stable, versioned APIs, and publish domain events outward into a durable log that downstream analytics, retrieval and agent workloads consume. The core stays clean; the history lives where it can be replayed without hammering production. Our SAP practice spends a lot of its time on exactly this boundary — deciding which events are worth publishing as contracts and which are implementation noise. Related patterns show up across our client work in commerce and service operations.

The test to run this quarter

Choose one business domain. Stand up a brand-new consumer. Replay ninety days of history into it, unaided by the source system, and compare the reconstructed state against production. Time how long it takes and count how many people you needed. That number — not your broker vendor, not your API gateway tier — is the honest measure of how ready your architecture is for AI workloads. Organisations that can replay confidently will keep adding intelligent consumers cheaply. Those that cannot will rebuild bespoke data plumbing for every new model, and will keep mistaking the cost of that plumbing for the cost of AI.

Frequently Asked Questions

Do we need event sourcing to get replay capability?

No. Full event sourcing — where the log is the system of record — is a heavier commitment than most enterprises need. Publishing durable, well-versioned domain events from systems that keep their own transactional state gives you replay for downstream consumers without rearchitecting the core.

How long should we retain events?

Set it per domain based on the longest legitimate replay need, typically driven by audit, model retraining or backfill scenarios rather than by operations. With tiered storage to object storage, months or years of retention is usually affordable; the harder constraints are privacy obligations and schema evolution, not cost.

Where do API gateways fit if events carry the history?

The gateway still governs the synchronous write and current-state read path — authentication, idempotency, quotas and versioning, increasingly for agent callers rather than human-built clients. Think of the gateway as governing intent and the event log as governing memory; both need catalogues and contracts.

Comments (0)

Leave a Comment

No comments yet. Be the first to comment!