Table of Contents
Table of Contents
A Kafka cluster can look healthy at its average write rate and still fall behind when traffic arrives in a burst. A flash sale, replay job, telemetry fan-in, or batch release can push producers above the rate that brokers, caches, disks, and network paths were sized to handle. The question for a diskless Kafka design is not whether object storage can hold the retained data. It is whether every part of the write and read path can absorb the burst, recover from a miss, and return to normal without permanent over-provisioning.
The practical way to answer that question is to turn “burst” into measurements. Record the baseline rate, peak rate, burst duration, record shape, consumer freshness objective, and the storage and network paths used at each phase. Then test the failure and cost boundaries before choosing an architecture. A shared-storage Kafka design can change the capacity problem, but it does not remove the need for a capacity model.
1What a burst workload means in a diskless Kafka design
A burst has three dimensions: how far the rate rises above baseline, how long it lasts, and what happens to consumers while it is happening. Two workloads with the same peak throughput can create different incidents if one has large records, many partitions, or a consumer group that must catch up from object storage. Treat peak throughput as an input, not as a capacity answer.
Start with a small worksheet for each topic or workload class. Let B be the sustained baseline rate, P the peak rate, and T the burst duration. The ingress volume added by the burst is (P - B) × T; the useful capacity question is where that volume waits before consumers process it. Depending on the implementation, it may be held in broker memory, a write-ahead log (WAL), a remote object store, or a consumer backlog. Keep the calculation in bytes and use the same unit across producer, broker, storage, and network measurements.
The worksheet should also capture partition distribution and record shape. A burst that raises total throughput evenly across partitions stresses shared bandwidth. A burst concentrated on a few partitions stresses leader CPU, request queues, cache admission, and the network path for those partitions. Compression changes the bytes crossing each boundary, while large records change fetch and object-request behavior. These details decide whether a capacity plan is safe.
Use the following questions to turn the worksheet into a decision:
- Ingress: What are the baseline and peak producer rates for each topic, and how are they distributed across partitions?
- Retention: Which bytes must be immediately durable, and which bytes can be served from a remote store after the burst?
- Freshness: How much consumer delay can the application tolerate while it catches up?
- Recovery: What must happen if a broker, cache, object-storage endpoint, or network path fails during the peak?
The output should be a set of limits and actions. “The cluster has enough storage” is not a limit; “the consumer must return to its freshness objective after the peak under a cold read” is one.
2The mechanism: brokers, cache, WAL, and object storage
Diskless Kafka changes where persistent stream data lives, so burst behavior is governed by a chain of buffers rather than a broker disk-size setting. The broker handles Kafka requests, partition leadership, and scheduling. A cache serves hot data. A WAL can provide a durable write boundary before data is uploaded and compacted into object storage. Object storage provides the durable retention layer and the catch-up read path. Each layer has a different queue, latency, and failure mode.
| Layer | What to measure during a burst | Failure question |
|---|---|---|
| Broker compute | Request queue time, CPU, partition leadership, and produce/fetch errors | Can the broker keep accepting work when a few partitions become hot? |
| Cache | Hit rate, eviction, memory pressure, and tailing versus replay reads | Does a replay displace hot data and raise consumer latency? |
| WAL | Append latency, capacity headroom, flush backlog, and recovery time | Where do acknowledged writes remain if upload pauses? |
| Object storage | Request latency, error and retry rate, bytes, and request volume | Can catch-up reads and background uploads share the endpoint safely? |
| Network | Client-to-broker and broker-to-storage bytes, route, and congestion | Does placement introduce a cost or a failure boundary during the peak? |
The cache and WAL are not interchangeable. A cache can be evicted because it is serving a different read pattern. A WAL is a durability boundary whose behavior must be tested when the remote store slows down. If the WAL fills while object-storage uploads are blocked, the system needs a defined backpressure or admission response. Measure that response instead of assuming the remote store will always drain the queue.
Read behavior deserves the same attention as writes. Tailing reads stay near the producer position and can be served by memory or the WAL. Catch-up reads start at older offsets and may require object-storage reads and prefetch. A burst followed by a consumer restart therefore exercises a different path from a burst followed by healthy tailing. Test both paths with the same partition and record distribution.
The network path is part of the storage model. Record whether clients, brokers, WAL storage, and object storage share a zone, cross zones, or use a private endpoint. A design can pass a throughput test while accumulating an unexpected transfer bill or a single route dependency. Keep network bytes and endpoint errors in the same run log as producer and consumer metrics.
3Failure, cost, and compatibility checks
A burst plan is production-ready when it describes what the system does under pressure, not only how it behaves on the happy path. Run a controlled test with a representative record shape and a bounded peak. During the test, watch the following gates:
- Admission gate. Producers receive the configured acknowledgement behavior, retries stay within the application policy, and partition hotspots are visible. If a producer is throttled, the run log should show when and why.
- Durability gate. Acknowledged records remain readable after an upload pause or broker restart. Record the WAL or storage boundary at which the write is considered durable for the selected implementation.
- Catch-up gate. A consumer group that falls behind can replay older offsets without evicting the data needed for tailing workloads. Measure freshness recovery with a warm cache and a cold cache.
- Recovery gate. Replace or restart a broker during the burst, then confirm metadata ownership, partition serving, and consumer progress. The recovery test must include any remote-store or network errors that operators could encounter.
- Cost gate. Separate compute, object-storage capacity, object-storage requests, WAL, observability, and network lines. Use the selected provider’s current price sheet and endpoint topology; do not fold them into a single “storage cost.”
Compatibility is another gate because a diskless architecture still has to carry Kafka application behavior. Check producer acknowledgements and idempotence, consumer commits and rebalances, transactions where used, compaction and retention policies, Kafka Connect tasks, Kafka Streams state handling, authentication, quotas, and monitoring integrations. A client that can connect is only the first test.
For an existing Kafka deployment, compare the burst test with the current storage model. Kafka Tiered Storage keeps an active local tier and moves older segments remotely; KIP-405 describes that direction. The diskless topics proposal is a useful reference for the Apache Kafka discussion, but a proposal does not establish the release, feature coverage, or operational behavior of the implementation you will run. Test the actual release and workload.
4How AutoMQ changes the operating model
Once the measurements show that broker-local retention and burst headroom are the main constraints, the architectural requirement becomes clearer: compute should be able to scale independently from retained stream data, while the write path still needs an explicit durability boundary. AutoMQ is a Kafka-compatible cloud-native streaming platform that uses a shared-storage architecture. Its brokers handle Kafka protocol and partition work while S3Stream writes durable stream data to S3-compatible object storage, with cache and WAL behavior selected for the deployment.
That separation changes the burst decision in two ways. First, adding or removing broker compute does not require copying a broker’s entire retained log directory to a replacement node. The team can evaluate compute headroom and storage headroom as separate measurements. Second, a WAL and cache make the hot write and read paths explicit. The chosen WAL type, object-storage endpoint, and cache policy still determine latency, cost, and failure behavior, so they belong in the same worksheet as the workload peak.
This is where a neutral test protects the design from marketing shorthand. Do not treat “diskless” as “remote storage handles every operation directly.” Verify where acknowledged writes land, which reads are cache hits, when object storage is involved, and what another broker can recover after a failure. AutoMQ’s diskless engine overview and S3Stream architecture documentation explain the storage path; your own burst and recovery runs determine whether it fits your service-level objectives (SLOs).
A shared-storage architecture also changes scaling operations. The scaling trigger can be broker CPU, request queue time, partition hotness, cache pressure, or consumer freshness rather than local-disk occupancy alone. Scaling policy should name the signal that starts an action and the signal that ends it. Otherwise a team may add brokers after the burst has already caused consumer lag, or keep extra brokers running after the workload returns to baseline.
5Decision checklist and FAQ
Before approving a diskless Kafka burst design, keep one signed record for the workload class. It should answer these questions:
- Are
B,P, andTmeasured for each topic class, with partition skew and record size included? - Which bytes are durable at producer acknowledgement, and how is that boundary verified during an upload pause?
- Can a consumer replay from older offsets without starving tailing reads?
- Which broker, cache, WAL, object-storage, and network signals trigger an operator action?
- Are compute, storage, request, WAL, observability, and network costs calculated from current provider inputs?
- Have the Kafka client features in scope been exercised during a burst and a broker recovery?
- Is the rollout reversible, with a known stop condition and an owner for the decision?
5.1Does diskless Kafka remove the need to plan for bursts?
No. It changes the storage and scaling boundaries. You still need a rate-and-duration model, partition distribution, cache and WAL measurements, object-storage request limits, network placement, and consumer recovery tests.
5.2Is a burst only a producer problem?
No. The producer peak is the visible trigger, but consumer catch-up, cache eviction, object-storage reads, and network paths often determine how long the system remains under pressure. Test the burst and the recovery window as one scenario.
5.3How should a team compare diskless Kafka with tiered storage?
Use the same workload and ask where active data, durable writes, replay reads, and partition ownership live. The compute-storage separation comparison can provide context, but the decision should come from your measurements and release-specific feature checks.
A burst workload is a capacity event followed by a recovery event. Size the first one with rate, duration, partition skew, and storage boundaries. Validate the second one with cold reads, failure drills, consumer freshness, and a cost record. If a shared-storage Kafka architecture passes those gates, start an AutoMQ evaluation using the same workload trace and stop conditions.
