Blog

Diskless Kafka Exactly Once: A Production Framework

Table of Contents

Table of Contents

“Exactly once” is often stated and hard to measure in a diskless Kafka deployment. A producer can retry after a timeout, a broker can fail after accepting a batch, a transaction coordinator can restart halfway through a commit, and a consumer can crash after applying a side effect but before committing its offset. Moving Kafka data to object storage changes where those events are persisted and replayed. It does not make them disappear.

For a diskless Kafka design, exactly-once semantics are a contract across three boundaries: the Kafka protocol, the storage path that preserves records and transaction state, and the application that consumes the result. The useful production question is therefore not “does Kafka on S3 support exactly once?” It is: can the team reproduce a failure, identify the last safe boundary, and prove that a replay does not create a second effect?

Exactly-once decision map showing protocol guarantees, storage-path measurements, failure tests, and rollout gates

1What exactly-once semantics cover in a diskless Kafka design

Kafka’s exactly-once story has several layers that are often collapsed into one phrase. An idempotent producer prevents retrying a request from appending a second record to the same partition when the original append succeeded but the response was lost. A transactional producer can group records written to multiple partitions with the consumer offsets that represent the work. A consumer configured with isolation.level=read_committed can hide records from aborted transactions. Those guarantees are described in the Apache Kafka semantics documentation, but each one has a defined boundary.

That boundary matters when the output leaves Kafka. A Kafka transaction can atomically publish output records and commit source offsets inside Kafka. It cannot automatically roll back a payment API call, a database write, or an email sent by a consumer that is outside the transaction. The downstream system must participate in the transaction or provide an idempotent operation keyed to the event. Treating Kafka’s transaction protocol as an end-to-end business guarantee is how a test plan becomes a production incident.

The storage architecture adds another boundary. A diskless broker can keep hot data in Data caching, persist the write path through WAL (Write-Ahead Log) storage, and place durable stream data in S3 storage. A consumer replay can therefore cross a cache boundary and an object-storage request even though the Kafka client still sees a normal fetch response. The semantics remain a Kafka contract; the evidence has to include the layers that make that contract durable.

Use this vocabulary before designing a test:

LayerContract to verifyWhat a failure can look like
ProducerIdempotent sequence and producer epoch remain valid across retry and session recoveryA timeout is followed by a duplicate or an out-of-order sequence error
Transaction coordinatorCommit, abort, and transaction timeout decisions are durable and visible to the right consumersA transaction remains pending, or a consumer sees an aborted record
Kafka storage pathRecords, transaction markers, offsets, and metadata survive the selected failureA replay stops at a missing range or reads a marker later than the data it fences
Application sinkThe external effect is atomic or idempotent with the Kafka progress it acknowledgesA consumer retry creates two rows, charges, or API requests

The table is a reminder that “exactly once” is a chain. The weakest untested boundary defines the result that customers experience.

2The mechanism: brokers, cache, metadata, and object storage

Traditional Apache Kafka uses a Shared Nothing architecture. A broker owns local log segments for its Partitions, and ISR (In-Sync Replicas) replication keeps copies on other brokers. The transaction protocol is implemented above that storage layer: producer IDs and epochs, transaction markers, group offsets, and the transaction state log must all be ordered and recovered consistently with the records they describe.

A diskless Kafka architecture replaces the broker-local persistence assumption with a Shared Storage architecture. The broker still handles Kafka requests and partition leadership, but durable stream data is served through a storage layer that may include a cache, WAL, and object storage. This changes the failure path without changing the client API. A broker replacement can fetch a historical range from S3 storage instead of rebuilding it from a local replica, while a hot write can be acknowledged from WAL storage before background upload completes, according to the chosen deployment and configuration.

Diskless Kafka data path showing producers, brokers, transaction metadata, WAL, Data caching, S3 storage, and consumers

Exactly-once tests need to follow both data and control signals. For a transactional read-process-write application, a representative path looks like this:

  1. The consumer reads source records with the intended isolation level and records the source Offset it has processed.
  2. The producer sends output records with an idempotent Producer identity and a transaction boundary that matches the processing unit.
  3. The broker appends records and transaction markers through the selected WAL and storage path, while KRaft metadata and the transaction state log track ownership and commit state.
  4. The producer commits the transaction, including the source offsets when the application uses Kafka’s consume-transform-produce pattern.
  5. A read_committed consumer reads the output and verifies that committed records are visible while aborted records remain hidden.

The critical observation is step three. The protocol does not care whether the bytes came from a local log, a WAL, or S3 storage, but recovery does. A cache miss during a replay can add object-storage reads. An object-storage error can delay the visibility of a transaction marker. A broker restart can expose a gap between records that reached WAL and records that were uploaded to S3. None of those observations proves a semantic violation by itself, yet each one belongs in the evidence timeline.

WAL storage is often misread. AutoMQ documents WAL as a fixed-size, cyclic persistence layer used for low-latency writes and recovery of data that has not reached S3 storage. It is not the long-term source of retained history. The WAL storage documentation describes the available WAL types and their different failure domains. A test should record which WAL type is active and what it is expected to recover; otherwise “diskless” becomes an imprecise label for several different durability paths.

The same distinction applies to Tiered Storage. KIP-405 describes a local tier and a remote tier, while a diskless design makes shared storage the durable boundary for the stream. Both designs can issue object-storage requests, but a transaction replay can encounter different ownership, cache, and recovery behavior. Verify the release and implementation instead of inferring semantics from the word “remote.”

3Failure, cost, and compatibility checks

An exactly-once test is useful only when its pass condition is written before the fault. Start with the application’s contract, then map the fault to the record, marker, Offset, and external effect that should be observable.

DrillEvidence to collectPass condition
Producer timeout after a transactional batch is acceptedProducer ID, sequence, retry, broker response, transaction state, and output record countThe retry does not append a second record, or the client receives a documented fencing or sequence error and the application recovers without a second effect
Broker failure before the transaction commit responseAppend acknowledgement, WAL status, transaction marker, coordinator epoch, and client retryAfter recovery, read_committed exposes only the committed result and the application can explain the fate of the in-flight transaction
Coordinator restart during commitTransaction state transitions, __transaction_state activity, producer epoch, and commit or abort responseThe transaction reaches one terminal state, and consumers do not observe a partial output
Cache miss or cold replay from S3 storageRequested Offset range, Data caching state, S3 reads, retries, and fetch latencyThe replay returns the same committed records and markers as the control run, with bounded and observable backpressure
Object-storage error while a transaction is openStorage error, WAL recovery, producer timeout, transaction timeout, and alert ownershipThe team can distinguish a storage stall from a protocol abort and has a documented retry or abort action
Consumer crash after processing and before offset commitApplication effect key, committed Offset, restart position, and downstream responseA replay either produces the same idempotent effect or is rejected by a deduplication key; the result is not assumed from Kafka alone
Consumer with read_uncommittedIsolation level, aborted-record markers, and application outputThe test demonstrates why this consumer is outside the exactly-once read contract and is not used for the production path

The no-fault control run is part of every row. Use the same keys, partition assignment, transaction size, client versions, storage configuration, and downstream system. Without that control, a slow sink or a pre-existing lag spike can be mistaken for a storage failure.

Keep the accounting separate from the semantic result. A replay can increase S3 request volume, WAL occupancy, cache churn, or network transfer while still returning the correct records. Capture those lines separately from broker compute and downstream costs. If an object-store request or cross-AZ (Availability Zone) route is part of the chosen deployment, record it as a cost and an operational dependency, not as evidence that exactly-once semantics are stronger or weaker.

Compatibility needs the same discipline. Inventory the producer and consumer client versions, transaction settings, authentication, ACLs, consumer isolation, Kafka Streams or Connect behavior, and the downstream system’s idempotency contract. The KIP-1150 Diskless Topics proposal is useful context for the diskless topic concept, but it is not proof that a particular Kafka release or managed service implements the proposal. Validate the exact release, client, storage mode, and failure drill together.

The most useful artifact is a single timeline that joins protocol events to storage events:

plaintext
T0  transactional batch begins
T1  records are appended and the client receives or loses the response
T2  fault is injected or detected
T3  coordinator and producer epochs change, or remain stable
T4  WAL and S3 storage report recovery or retry state
T5  transaction reaches commit or abort
T6  source Offset is committed, or the application restarts before that commit
T7  read_committed replay verifies output records and markers
T8  downstream effect is reconciled with the event key

The labels are a measurement model, not a promise about event ordering in every implementation. Map each label to a log, metric, trace, or client observation before the run. If the timeline cannot explain why a replay was safe, the production gate is still open.

4How AutoMQ changes the operating model

The worksheet first asks which Kafka semantics must remain true and then asks where the durable bytes and transaction metadata live. That leads naturally to a Kafka-compatible Shared Storage architecture, where brokers handle the protocol while shared storage carries the stream. AutoMQ is one such cloud-native streaming platform. Its compatibility documentation describes reuse of the Apache Kafka computing layer and compatibility with Kafka clients and ecosystem components; the architecture overview explains how AutoMQ Brokers, S3Stream, WAL storage, Data caching, and S3 storage fit together.

That architecture changes the evidence you need from a broker replacement. The replacement does not begin by copying a complete retained log from a broker-local disk. It must regain the expected metadata ownership, reach the configured WAL and S3 storage, recover any data that is still in the WAL path, and serve committed and historical reads. The transaction protocol still needs producer fencing, transaction-state recovery, marker ordering, and offset continuity. Shared storage changes where the records are recovered; it does not lower the standard for proving a committed transaction.

WAL choice must stay visible in the pilot. AutoMQ Open Source supports S3 WAL, while AutoMQ commercial editions can offer other WAL types depending on deployment. The selected type changes latency, topology, and the set of failures that the test can absorb. Record the WAL type, S3-compatible object-storage endpoint, network boundaries, cache policy, and metadata configuration next to every result. A passing test for one storage path is evidence for that path and workload.

Kafka compatibility is also a workload claim. AutoMQ’s compatibility documentation describes the intended boundary, but a production team should still test its actual producer libraries, transaction IDs, consumer isolation, Kafka Streams state stores, Connect plugins, and administrative tooling. A protocol check cannot reveal an application that applies a duplicate external effect after a restart.

Two related checks help fill out the same worksheet: the Kafka compatibility guide focuses on client and ecosystem behavior, while producer acknowledgment policies explain why an acknowledgement setting is only meaningful when its durability boundary is understood. Neither replaces the failure drills here; both make the test inputs more explicit.

The migration path should preserve the worksheet. Move one production-shaped transactional flow, keep its source Offset and downstream effect evidence, and repeat the producer retry, coordinator restart, broker replacement, cold replay, and sink recovery drills. Compare the same record keys and transaction boundaries. If shared storage removes local-log movement but adds object-store request pressure, the result should show the assigned owner and the action it requires. The goal is a measurable operating model, not a storage label.

5Decision checklist and FAQ

Production readiness scorecard for exactly-once Diskless Kafka: protocol, storage, recovery, application, observability, and rollout gates

Before expanding a Diskless Kafka pilot, require evidence for each gate:

  • Protocol gate: Idempotent Producer retries, transaction commit and abort, producer fencing, transaction timeouts, and read_committed behavior pass with the client versions used in production.
  • Storage gate: The selected WAL type, Data caching behavior, S3 storage access, metadata recovery, and object-storage failure response are visible in the same timeline as the Kafka events.
  • Replay gate: A broker or coordinator restart reproduces the expected records, markers, Offsets, and transaction terminal state without relying on an operator’s memory.
  • Application gate: Every external side effect has transaction participation, an idempotency key, or a documented reconciliation procedure. Kafka’s commit alone is not treated as proof.
  • Operations gate: Alerts identify whether the owner is the Producer, Consumer, coordinator, WAL storage, S3 storage, network, or downstream system, and the runbook names the next action.
  • Rollback gate: The team has a stop condition, a source Offset strategy, a downstream reconciliation step, and a tested way to return to the previous deployment.

5.1Does Diskless Kafka change Kafka’s exactly-once API?

The client still uses Kafka’s Producer, transaction, Offset, and Consumer isolation APIs. The storage path changes how records, markers, and metadata are persisted and recovered. API compatibility therefore needs to be paired with failure and replay tests.

5.2Does S3 storage make a transaction atomic?

No. S3 storage is a persistence layer in the data path. Kafka transaction atomicity still depends on the broker and coordinator protocol, and an external system still needs transaction participation or idempotent writes for an end-to-end business guarantee.

5.3What should a read_committed test prove?

It should show that committed output records become visible together with the intended Offset boundary and that aborted records remain hidden. Repeat the check after a producer retry, coordinator restart, broker replacement, and cold replay so the result is tied to recovery behavior.

5.4Is a WAL acknowledgement the same as a committed transaction?

No. WAL acknowledgement describes the durability boundary for a write path. A Kafka transaction remains pending until the coordinator records a commit or abort, and consumers must apply the configured isolation level. Keep WAL state and transaction state as separate evidence fields.

5.5Does Kafka exactly once cover a database or API call?

Only when the downstream operation participates in the same transaction boundary or is idempotent and reconciled with the Kafka Offset. Otherwise, a consumer crash between the side effect and the Offset commit can cause a replay.

5.6Where does AutoMQ fit?

AutoMQ is a Kafka-compatible cloud-native streaming platform built on Shared Storage architecture. Evaluate it with the same worksheet: the selected WAL type, S3-compatible storage, metadata and fencing behavior, client compatibility, cold reads, and downstream effect handling. The architecture changes which work a broker replacement performs; the test determines whether the deployment meets its exactly-once contract.

The useful result is not a green transaction metric. It is a timeline that explains which Producer attempt was accepted, which transaction reached a terminal state, which storage layer served the replay, and whether the downstream effect was applied once. Run that timeline against the current cluster and a Kafka-compatible Shared Storage candidate. When the evidence is ready, start an AutoMQ evaluation with the failure matrix and rollback conditions in hand.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.