Table of Contents
Table of Contents
A diskless Kafka design does not remove capacity planning. It changes the question from “How many disks do the brokers need?” to “Where do the bytes, requests, and recovery work go under the workload we actually run?” That shift matters before a migration, because a cluster can have ample object-storage space and still fail its latency or recovery objective when cache, WAL, request, or network capacity is missing.
The useful unit is a measured path. Producer bytes enter through brokers, recent data sits in memory or a write buffer, durable data moves to object storage, and consumers may read both hot and cold ranges. Metadata and access control follow a separate control path. A production plan has to account for each path, assign an owner, and record the evidence that sets its limit.
1What capacity planning means in a diskless Kafka design
Start with the workload contract rather than an instance type. Record producer ingress, consumer egress, retention and replay windows, partition distribution, compression behavior, peak shape, and the latency objectives that applications actually enforce. Use observations from the existing cluster where they exist; when a value is unknown, keep it as an explicit assumption and give it a date for validation.
A compact worksheet can express the storage side without pretending that every workload behaves alike:
- Ingress bytes per day = observed average producer rate × seconds in a day. Keep peak rate beside the average; the peak drives buffers and scale-out even when retention follows the average.
- Retained bytes = ingress bytes per day × the retention window, adjusted with the measured compression and object-layout overhead for the chosen implementation.
- Read demand = consumer fetch traffic plus catch-up and replay traffic. Fan-out makes this different from write volume, so do not infer it from producer metrics alone.
- Recovery demand = the bytes and metadata that a replacement broker must read before it can serve the workload at the required level.
These are relationships, not promises. Measure each term with the same unit and time window, then compare the result with service thresholds. If a team cannot explain where a number came from, that number is a planning risk rather than a capacity input.
The plan should also separate four questions that local-disk Kafka often hides inside one broker-size decision:
| Planning area | Evidence to collect | Capacity decision |
|---|---|---|
| Compute | Request rate, active partitions, protocol operations, CPU, and memory | Broker count and instance shape for the target workload |
| Cache and WAL | Hot-read ratio, write queue, flush time, cache evictions, and backpressure | Memory and durable write-buffer headroom |
| Object storage | Retained bytes, upload and fetch rates, object count, request rate, and lifecycle policy | Bucket, request, and service-limit planning |
| Network | Producer, consumer, object-store, management, and cross-zone paths | Bandwidth, endpoint, and placement decisions |
The table is deliberately broader than “storage size.” A diskless cluster can run out of useful cache, object-store request capacity, or network headroom while its bucket has plenty of free space. Capacity planning is complete only when the slowest path has a threshold and an operator who can act on it.
2The mechanism: brokers, cache, metadata, and object storage
A diskless architecture separates durable stream data from the broker that serves it. Brokers still own protocol handling, partition leadership, quotas, and request scheduling. They do not need to be the long-term owner of every retained byte. That distinction lets compute respond to traffic while retention follows the storage layer, but only when the intermediate paths are sized for the workload.
The write path usually has three observable stages. A broker accepts a produce request and appends data to a durable write buffer or WAL. The system acknowledges according to its durability contract, then uploads and organizes data in object storage. The plan therefore needs a write-acknowledgement objective, WAL flush capacity, upload bandwidth, and an alert for data that remains in the buffer longer than the recovery policy allows.
The read path has a different shape. Recent tailing reads may be served from memory or data close to the write path, while catch-up reads pull older ranges from object storage and place them in a cache. A replay-heavy consumer can therefore compete with normal traffic for network, object-store requests, and cache space. Measure those classes separately instead of hiding them in an average fetch rate.
Metadata is another capacity boundary. Topics, partitions, ownership, offsets, object indexes, ACLs, and controller traffic do not have the same byte profile as message data, but they determine whether a broker can route and recover that data. Track metadata growth, controller request latency, and the time required to rebuild ownership information during a node or zone event.
A useful plan names the failure signal for each component:
| Component | What to measure | What an undersized component looks like |
|---|---|---|
| Broker compute | CPU, request latency, active partitions, network throughput | Client latency rises while storage paths remain healthy |
| Data cache | Hit ratio, evictions, memory pressure, cold-read latency | Replay causes tail latency or repeated object reads |
| WAL or write buffer | Flush latency, queue depth, upload lag, backpressure | Produce acknowledgements slow or writes are throttled |
| Object storage | Read/write latency, request rate, errors, throttling, retained bytes | Upload backlog, catch-up delay, or retry storms appear |
| Metadata and control | Controller latency, quorum health, metadata size, recovery time | Ownership changes or failover take longer than the gate |
| Network boundaries | Bytes and latency by zone, endpoint, VPC, and region | The bill or the SLO changes when placement changes |
This is where “Kafka on S3” needs careful language. Object storage can hold the durable stream, but it does not make every request path equivalent. Request rate, object size, cache policy, endpoint placement, and the chosen write buffer influence both cost and latency. A plan that records only bucket growth is missing the operating model.
Diskless Kafka is also different from Kafka Tiered Storage. KIP-405 describes moving older log data to a remote tier while brokers continue to manage an active local tier. A fully shared-storage design changes the ownership boundary for durable data rather than treating remote storage as an extension of broker disks. KIP-1150 is a useful reference for the diskless-topic problem space and its semantics; verify the proposal’s status and supported behavior against the Kafka release you are evaluating.
3Failure, cost, and compatibility checks
Capacity numbers are meaningful only against failure scenarios. Ask what happens when a broker disappears during a write burst, when object storage returns errors, when a consumer starts a long replay, or when a zone loses connectivity. For each scenario, write the expected signal, the data that must be recovered, the owner who responds, and the point at which the workload should be throttled or moved.
A neutral evaluation worksheet can look like this:
| Scenario | Measurement | Production gate |
|---|---|---|
| Broker replacement | Time to restore request service, replay scope, and client error rate | Recovery fits the stated RTO and does not rely on an unmeasured cache assumption |
| Object-store slowdown | Upload lag, fetch latency, retries, and request errors | Backpressure and alerting protect acknowledged data and expose the degraded path |
| Replay surge | Consumer fetch rate, cache occupancy, object-store bandwidth, and application lag | Replay remains within the latency and cost envelope for the workload class |
| Zone or endpoint issue | Traffic by boundary, failed requests, and failover behavior | Placement and access paths have a tested fallback |
| Metadata pressure | Controller latency, metadata size, and ownership-change time | Recovery and scale operations stay within the operational window |
Cloud cost follows these paths too. Keep compute, object storage, requests, data transfer, observability, and engineering operations as separate lines in the model. Cross-zone replication, cross-region movement, private endpoints, and egress can have different pricing rules; use the provider’s current pricing page for the regions and paths in scope instead of carrying a remembered rate into the forecast.
Compatibility is a capacity input because an unsupported client behavior creates operational work. Inventory client versions, producer acknowledgements, idempotence and transactions, compacted topics, consumer-group behavior, Kafka Streams, Connect, admin tools, authentication, quotas, and monitoring integrations. A design that saves storage but forces a rewrite of a high-volume producer has changed the project being priced.
Migration capacity belongs in the same worksheet. Record the topics that move first, the read and write paths during cutover, the offset or replay strategy, the rollback point, and the extra capacity needed while both systems are active. “The bucket is ready” is not a migration gate; the application contract and the recovery path must be ready too.
4How AutoMQ changes the operating model
Once the neutral framework is written down, a Kafka-compatible shared-storage platform becomes a concrete option to test. AutoMQ keeps the Kafka protocol and semantics at the client boundary while using a Shared Storage architecture for durable stream data. Its S3Stream storage library, WAL layer, data cache, and S3-compatible object storage form a path that can be measured with the same worksheet above.
The architecture changes what a broker-size decision means. Compute can be sized for request and partition work while retained bytes live in shared storage. Adding or replacing a broker does not require treating its local disk as the source of truth for every partition. That can reduce data-movement pressure during scaling, but the claim still needs a workload test: measure request capacity, cache behavior, WAL flushes, object-store traffic, and recovery time under the conditions that matter to your team.
WAL choice is part of the plan, not a footnote. AutoMQ documents S3 WAL, EBS WAL, Regional EBS WAL, and NFS WAL as storage options with different latency, infrastructure, and failure-domain implications. The WAL storage documentation should be read alongside the deployment model and the target workload. Record the selected backend, its zone behavior, its recovery path, and the measurements that justify its cache and bandwidth settings.
The object store is still a production dependency. Plan bucket lifecycle and retention separately from broker scaling, watch upload and fetch rates, and validate request limits and access policies in the selected cloud or S3-compatible service. Network locality matters as well: measure the path from clients to brokers and from brokers to storage, then verify the zones, endpoints, IAM permissions, and egress rules that make the path real.
This is the useful distinction between a product claim and an operating model. AutoMQ can provide the Kafka-compatible Shared Storage architecture; the platform team still owns the thresholds, dashboards, failure drills, and rollout gates. The evidence should tell you whether the architecture fits your workload before a larger migration makes the answer expensive to change.
5Decision checklist and FAQ
Use this checklist at a design review, before a pilot, and again before production traffic moves:
- Workload contract: ingress, egress, retention, replay, partition distribution, compression, and peak shape have owners and measurement windows.
- Byte-path model: broker, cache, WAL, object storage, metadata, and network limits are separate lines with units and thresholds.
- Failure evidence: broker replacement, object-store degradation, replay surge, zone loss, and metadata recovery have been tested or explicitly scheduled.
- Compatibility inventory: clients, transactions, compaction, Consumer groups, Connect, Streams, admin tools, authentication, quotas, and metrics are mapped.
- Cost model: compute, storage, requests, transfer, observability, and migration operations use current provider inputs and stated assumptions.
- Rollout gate: the pilot has a stop condition, rollback point, dashboard owner, and review date.
5.1Does diskless Kafka mean the cluster has unlimited capacity?
No. Object storage can provide a different scaling boundary for retained bytes, but brokers, caches, WAL, object-store requests, metadata, and network paths still have limits. Capacity planning moves those limits into a visible worksheet.
5.2Is diskless Kafka the same as Kafka Tiered Storage?
No. Tiered Storage commonly keeps an active local tier and moves older segments to remote storage. A shared-storage design makes durable stream data independent of a broker’s local disk. Confirm the exact semantics and supported features of the implementation under evaluation.
5.3What should be measured before a pilot?
Measure the workload contract and the byte paths: producer and consumer rates, peak shape, retention, replay, cache behavior, write-buffer lag, object-store latency and requests, metadata recovery, network boundaries, and client compatibility. The measurements should match the failure and rollout gates you intend to use in production.
5.4How should AutoMQ appear in the capacity model?
Model AutoMQ as a Kafka-compatible Shared Storage option with an explicit WAL backend, cache policy, object-store configuration, network placement, and recovery procedure. Compare those inputs with the same workload and governance requirements used for the current Kafka design.
The original question was “How much capacity does a diskless Kafka cluster need?” The answer is a traceable set of byte paths, failure limits, and owners. If your current plan still makes broker count carry the weight of retained data, replay, and recovery, run this worksheet against one production-shaped workload. Then start an AutoMQ evaluation with the measurements and stop criteria in hand.
