Table of Contents
Table of Contents
A broker restart is a useful smoke test. It is not a recovery test for a diskless Apache Kafka design.
When durable stream data moves away from broker-local disks, the failure path changes with it. A replacement broker must regain metadata ownership, reach the storage layer, recover any write-ahead data, rebuild the cache it needs, and convince clients to refresh their view of the cluster. A test that records only “the process came back” can miss a producer acknowledgement gap, a stale leader, a blocked object-storage request, or a consumer that cannot replay data that is no longer hot.
The practical question is therefore: what evidence proves that a diskless Kafka deployment has recovered? The answer is a timed chain from the injected fault to restored produce, fetch, and replay behavior. This framework turns that chain into test cases, measurements, and decision gates.
1Start with a recovery contract, not a failure injection
Recovery testing is only useful when the team has written down what “recovered” means. An RPO (Recovery Point Objective) describes the acknowledged data the service must preserve. An RTO (Recovery Time Objective) describes the time allowed before the service returns to its required level of operation. Neither objective is a property of “diskless” by itself. The selected WAL storage, object storage, controller quorum, network path, client retry policy, and workload determine what the test can prove.
Write the contract as observable conditions. For example, a test may require that records acknowledged before the fault remain readable, that records produced after ownership converges can be acknowledged, and that a consumer group can resume from a known Offset without duplicate or missing application-visible Records beyond its documented delivery semantics. Add the owner of each signal and the exact timestamp source. If a condition cannot be observed, it is a hope, not a test assertion.
Keep the normal workload running while the test injects one controlled fault. A quiet cluster can hide cache misses, backpressure, partition skew, and retry storms. The workload does not need to be production-sized, but it should contain the read and write shapes that drive the recovery path: tailing consumers, at least one Catch-up Read, a consumer group with committed Offsets, and records that are acknowledged immediately before the fault.
The Apache Kafka replication documentation is a useful baseline for the traditional Shared Nothing architecture. A diskless test should make the differences explicit instead of assuming that a familiar replica check covers every dependency.
2What a diskless recovery path actually contains
A diskless Kafka deployment still has state. The difference is where authoritative durable stream data lives and which state a replacement broker must rebuild. Treat these components as separate checkpoints:
| Checkpoint | What the test must establish | Evidence to capture |
|---|---|---|
| Broker compute | The failed process is fenced or unavailable, and a replacement can start with the expected identity and listeners. | Process and node events, listener health, client metadata responses |
| Metadata and ownership | The Active Controller has made a replacement ownership or leadership decision and stale operations cannot continue writing. | Controller events, leader epochs, ownership records, fencing signals |
| WAL storage | Data accepted at the write boundary is recoverable according to the chosen WAL type and deployment topology. | WAL append, upload, recovery, and reset events; acknowledged record checks |
| S3 storage | Durable stream data and metadata objects remain reachable with the required permissions and network path. | Object-storage request status, error rate, latency, and range reads |
| Data caching | Tailing reads and historical Catch-up Reads return records after cache loss or cold start. | Cache miss rate, prefetch activity, fetch latency, replay checksum |
| Kafka clients | Producers and Consumers refresh metadata, retry within their contract, and resume the expected workload. | Producer errors, consumer assignments, committed Offsets, lag, duplicates |
The order matters. A successful process start does not prove that ownership converged. Ownership convergence does not prove that a WAL record was uploaded or that a historical fetch works. A consumer that has resumed tailing may still fail when asked to rebuild state from an older Offset. The test should close each checkpoint before it declares the next one healthy.
This is also where Tiered Storage must be separated from a Diskless architecture. Kafka's KIP-405 describes remote storage for older log segments while the active log and broker-local replica model remain part of the system. A diskless design changes the durability boundary for the stream itself. Its recovery test therefore includes the shared storage and WAL path as primary dependencies, not only as a cold-data tier.
3Build a failure matrix that exercises the boundary
Do not begin with a large chaos campaign. Start with a matrix that maps one fault to one hypothesis, one observation window, and one pass condition. Keep the injection reversible and record the exact start and stop times in a monotonic clock as well as the platform's wall-clock logs.
| Fault | Hypothesis | Measurements | Pass condition |
|---|---|---|---|
| Broker process or node loss | Another broker can take the required ownership without copying a broker-local log. | Fault detection, controller decision, leader readiness, producer errors, consumer resumption. | The stated RPO and RTO contract is met, and the original acknowledged Records remain readable. |
| Cache eviction or replacement restart | A cold path can serve both current and historical reads. | Cache misses, prefetch, Catch-up Read latency, replayed Record checksum. | Tailing and replay workloads meet their separate service objectives. |
| WAL storage interruption | The write boundary fails in a visible and bounded way. | Producer acknowledgement behavior, WAL backlog, recovery and reset events. | No acknowledged data violates the RPO; the runbook identifies the next operator action. |
| Object-storage denial or elevated errors | The primary durability path is observable before it becomes a silent data gap. | Request errors, retries, upload backlog, fetch failures, alerts. | The platform applies the documented backpressure or failure behavior and the test restores cleanly. |
| Controller or metadata disruption | Ownership and fencing prevent stale writers after a coordination fault. | Controller quorum, leader epochs, fencing errors, stale operation count. | One owner serves the Partition and clients converge without split-brain writes. |
| Network isolation | The deployment's fault domain and client route match its design. | Per-path reachability, cross-zone traffic, client retries, storage calls. | The test shows which operations continue, which stop, and how the service returns. |
The matrix should include a no-fault control run. Without it, a slow producer, an already high consumer lag, or an object-storage rate limit can be mistaken for recovery damage. Run the control with the same record keys, partition shape, retention settings, and client configurations that the fault test uses.
An existing failure-injection exercise for Kafka offers a useful operational habit: write the hypothesis and the pass line before the injection. For a diskless cluster, add the storage boundary and WAL state to those fields. Otherwise, a test can finish with green broker health while its durable path is still stalled.
4Measure a timeline that an operator can replay
Recovery time is not one timestamp. Capture a timeline whose events map to an owner and a decision. A minimal event stream looks like this:
T0 fault injected
T1 failure detected and client symptoms begin
T2 failed Broker fenced or removed from ownership
T3 Controller records a replacement owner or leader
T4 replacement reaches storage and WAL
T5 recovered data is committed or made readable
T6 produce succeeds at the required acknowledgement boundary
T7 tailing Consumer resumes
T8 Catch-up Read and state rebuild pass
T9 recovery gate closed and evidence archived
The labels are a measurement model, not a promise about implementation details. A specific system may combine events or emit them in another order. The test owner should map each label to an actual log, metric, trace, or client observation before the run.
Measure separately for producers and consumers. Producer recovery can look healthy while a Consumer group remains stuck on a stale assignment or cannot read the historical range it needs. Record the committed Offset before the fault, the first successful fetch after recovery, and the application-level checksum or key set for a bounded replay window. Treat duplicates according to the producer and consumer delivery contract; do not call every retry a data loss event.
The recovery timeline also exposes hidden cost and capacity work. Object-storage requests may be retried while WAL data is uploaded. A cold cache may pull a large historical range while producers resume. A cross-zone route may appear during a node replacement even when the steady-state design keeps clients local. Keep those observations in the report, but do not convert one test run into a universal price or performance claim.
5Turn test results into rollout gates
A recovery test should end in a decision, not a dashboard screenshot. Use gates that a release manager and an on-call engineer can both understand:
- Data gate: every acknowledged Record in the test window is accounted for at the documented durability boundary, and replay verification has a reproducible result.
- Ownership gate: the failed Broker cannot continue writing, one current owner serves each tested Partition, and fencing or epoch evidence is archived.
- Client gate: producer retries, Consumer group assignment, committed Offsets, and application-visible delivery behavior match the contract.
- Dependency gate: WAL storage, S3 storage, metadata quorum, credentials, and network paths expose useful alerts and have a named owner.
- Operations gate: an operator can execute the runbook from the evidence collected during the drill, including rollback or escalation when the pass line is missed.
If one gate fails, preserve the workload and logs before resetting the environment. Recovery failures can be caused by cloud permissions, volume topology, stale metadata, or storage throttling rather than the Kafka request path itself. Treat WAL recovery, upload, reset, and stream close as separate evidence points, and check the current product documentation and deployment code when the path differs between versions or deployment modes.
Run the matrix again after changing the WAL type, object-storage endpoint, network topology, Kafka client version, retention policy, or controller configuration. Those changes can alter the recovery boundary even when the application code stays the same. A passing drill is evidence for the tested configuration and workload envelope.
6How AutoMQ changes the test shape
The test framework first asks where durable stream data lives, how writes become durable, and which signals show that a replacement is serving correctly. Those questions lead naturally to architectures that separate broker compute from durable storage while preserving the Kafka client contract. This is where AutoMQ, a Kafka-compatible cloud-native streaming platform, is a concrete architecture to evaluate.
AutoMQ's Shared Storage architecture uses S3Stream to place durable stream data in S3-compatible object storage, with WAL storage supporting the write and recovery path and Data caching supporting reads. The architecture overview describes the relationship between AutoMQ Brokers, KRaft metadata, S3Stream, WAL storage, and object storage. The WAL storage documentation is the place to confirm how a selected WAL type changes latency, topology, and recovery dependencies.
That architecture changes what the recovery test should look for. A Broker replacement does not start by rebuilding the complete retained log from a broker-local disk. It must regain the expected identity and metadata ownership, reach the configured storage, recover any data that remains in the WAL path, and serve both current and historical reads. The S3Stream overview gives the storage-layer context; the actual pass line still comes from the workload-specific drill.
Kafka compatibility belongs in the test matrix as well. Verify the client APIs, delivery semantics, Consumer groups, transactions, compaction, Connect, and administrative tools that the workload uses. AutoMQ's compatibility documentation documents the intended compatibility boundary. It does not replace a test of your exact client versions, plugins, retry settings, and replay code.
An uptime-claim evaluation checklist is useful here because it keeps architecture, evidence, and service commitments separate. A Shared Storage design can remove broker-local data movement from part of a replacement path. It cannot make an object-storage permission error, an unavailable metadata quorum, or an untested client recovery behavior disappear.
7Decision checklist and FAQ
Before calling a Diskless Kafka deployment ready for production, ask:
- Which Records are acknowledged before the fault, and where is that durability boundary recorded?
- Which WAL type and object-storage service are in the test, and what failure domain do they introduce?
- What proves that a stale Broker is fenced before a replacement serves writes?
- Which controller, metadata, storage, and client events delimit the recovery timeline?
- Does the test include a cold Catch-up Read as well as a tailing Consumer?
- What happens when object storage is slow or unavailable while the Broker process is healthy?
- Can an operator reproduce the pass or fail decision from archived logs, metrics, traces, and client output?
- Which changes require the matrix to run again before rollout?
7.1Is a Broker restart enough to test Diskless Kafka recovery?
No. It checks process startup. A production recovery test must also verify metadata ownership, fencing, WAL behavior, object-storage access, cache or cold-read behavior, and Kafka client recovery.
7.2Does diskless mean the service has no local state?
No. A Diskless architecture can still use caches, WAL storage, metadata, connections, and in-flight requests. The test must name which state is authoritative and which state is rebuilt after replacement.
7.3Should a recovery test use only tailing Consumers?
No. Tailing validates the hot path. Add a Catch-up Read or a bounded state rebuild so the test exercises historical data that may not be in memory after a replacement.
7.4Does Tiered Storage provide the same recovery boundary?
Not automatically. Tiered Storage can move older segments to remote storage while the active log and broker-local replica model remain. Read the implementation and test the actual durability and ownership paths instead of inferring them from the word “remote.”
7.5Where does AutoMQ fit?
AutoMQ is a Kafka-compatible platform built on Shared Storage architecture. Evaluate it by the same matrix: selected WAL type, S3-compatible storage, metadata and fencing behavior, client compatibility, cold reads, and operator evidence. The architecture changes which work a replacement must do; the test determines whether the deployment meets its recovery contract.
The useful result of a recovery drill is not a green Broker icon. It is a timeline that explains when acknowledged data remained durable, when ownership changed, when clients resumed, and which dependency still needs work. Start with one representative workload, archive that evidence, and repeat the matrix after each storage or topology change. When you are ready to run the same test against a Kafka-compatible Shared Storage deployment, start an AutoMQ evaluation with the failure matrix and recovery contract in hand.
