Blog

Diskless Kafka Capacity Planning: A Production Framework

Table of Contents

Table of Contents

A diskless Kafka design does not remove capacity planning. It changes the question from “How many disks do the brokers need?” to “Where do the bytes, requests, and recovery work go under the workload we actually run?” That shift matters before a migration, because a cluster can have ample object-storage space and still fail its latency or recovery objective when cache, WAL, request, or network capacity is missing.

The useful unit is a measured path. Producer bytes enter through brokers, recent data sits in memory or a write buffer, durable data moves to object storage, and consumers may read both hot and cold ranges. Metadata and access control follow a separate control path. A production plan has to account for each path, assign an owner, and record the evidence that sets its limit.

Diskless Kafka capacity planning decision map

1What capacity planning means in a diskless Kafka design

Start with the workload contract rather than an instance type. Record producer ingress, consumer egress, retention and replay windows, partition distribution, compression behavior, peak shape, and the latency objectives that applications actually enforce. Use observations from the existing cluster where they exist; when a value is unknown, keep it as an explicit assumption and give it a date for validation.

A compact worksheet can express the storage side without pretending that every workload behaves alike:

  • Ingress bytes per day = observed average producer rate × seconds in a day. Keep peak rate beside the average; the peak drives buffers and scale-out even when retention follows the average.
  • Retained bytes = ingress bytes per day × the retention window, adjusted with the measured compression and object-layout overhead for the chosen implementation.
  • Read demand = consumer fetch traffic plus catch-up and replay traffic. Fan-out makes this different from write volume, so do not infer it from producer metrics alone.
  • Recovery demand = the bytes and metadata that a replacement broker must read before it can serve the workload at the required level.

These are relationships, not promises. Measure each term with the same unit and time window, then compare the result with service thresholds. If a team cannot explain where a number came from, that number is a planning risk rather than a capacity input.

The plan should also separate four questions that local-disk Kafka often hides inside one broker-size decision:

Planning areaEvidence to collectCapacity decision
ComputeRequest rate, active partitions, protocol operations, CPU, and memoryBroker count and instance shape for the target workload
Cache and WALHot-read ratio, write queue, flush time, cache evictions, and backpressureMemory and durable write-buffer headroom
Object storageRetained bytes, upload and fetch rates, object count, request rate, and lifecycle policyBucket, request, and service-limit planning
NetworkProducer, consumer, object-store, management, and cross-zone pathsBandwidth, endpoint, and placement decisions

The table is deliberately broader than “storage size.” A diskless cluster can run out of useful cache, object-store request capacity, or network headroom while its bucket has plenty of free space. Capacity planning is complete only when the slowest path has a threshold and an operator who can act on it.

2The mechanism: brokers, cache, metadata, and object storage

A diskless architecture separates durable stream data from the broker that serves it. Brokers still own protocol handling, partition leadership, quotas, and request scheduling. They do not need to be the long-term owner of every retained byte. That distinction lets compute respond to traffic while retention follows the storage layer, but only when the intermediate paths are sized for the workload.

The write path usually has three observable stages. A broker accepts a produce request and appends data to a durable write buffer or WAL. The system acknowledges according to its durability contract, then uploads and organizes data in object storage. The plan therefore needs a write-acknowledgement objective, WAL flush capacity, upload bandwidth, and an alert for data that remains in the buffer longer than the recovery policy allows.

The read path has a different shape. Recent tailing reads may be served from memory or data close to the write path, while catch-up reads pull older ranges from object storage and place them in a cache. A replay-heavy consumer can therefore compete with normal traffic for network, object-store requests, and cache space. Measure those classes separately instead of hiding them in an average fetch rate.

Metadata is another capacity boundary. Topics, partitions, ownership, offsets, object indexes, ACLs, and controller traffic do not have the same byte profile as message data, but they determine whether a broker can route and recover that data. Track metadata growth, controller request latency, and the time required to rebuild ownership information during a node or zone event.

Diskless Kafka data path

A useful plan names the failure signal for each component:

ComponentWhat to measureWhat an undersized component looks like
Broker computeCPU, request latency, active partitions, network throughputClient latency rises while storage paths remain healthy
Data cacheHit ratio, evictions, memory pressure, cold-read latencyReplay causes tail latency or repeated object reads
WAL or write bufferFlush latency, queue depth, upload lag, backpressureProduce acknowledgements slow or writes are throttled
Object storageRead/write latency, request rate, errors, throttling, retained bytesUpload backlog, catch-up delay, or retry storms appear
Metadata and controlController latency, quorum health, metadata size, recovery timeOwnership changes or failover take longer than the gate
Network boundariesBytes and latency by zone, endpoint, VPC, and regionThe bill or the SLO changes when placement changes

This is where “Kafka on S3” needs careful language. Object storage can hold the durable stream, but it does not make every request path equivalent. Request rate, object size, cache policy, endpoint placement, and the chosen write buffer influence both cost and latency. A plan that records only bucket growth is missing the operating model.

Diskless Kafka is also different from Kafka Tiered Storage. KIP-405 describes moving older log data to a remote tier while brokers continue to manage an active local tier. A fully shared-storage design changes the ownership boundary for durable data rather than treating remote storage as an extension of broker disks. KIP-1150 is a useful reference for the diskless-topic problem space and its semantics; verify the proposal’s status and supported behavior against the Kafka release you are evaluating.

3Failure, cost, and compatibility checks

Capacity numbers are meaningful only against failure scenarios. Ask what happens when a broker disappears during a write burst, when object storage returns errors, when a consumer starts a long replay, or when a zone loses connectivity. For each scenario, write the expected signal, the data that must be recovered, the owner who responds, and the point at which the workload should be throttled or moved.

A neutral evaluation worksheet can look like this:

ScenarioMeasurementProduction gate
Broker replacementTime to restore request service, replay scope, and client error rateRecovery fits the stated RTO and does not rely on an unmeasured cache assumption
Object-store slowdownUpload lag, fetch latency, retries, and request errorsBackpressure and alerting protect acknowledged data and expose the degraded path
Replay surgeConsumer fetch rate, cache occupancy, object-store bandwidth, and application lagReplay remains within the latency and cost envelope for the workload class
Zone or endpoint issueTraffic by boundary, failed requests, and failover behaviorPlacement and access paths have a tested fallback
Metadata pressureController latency, metadata size, and ownership-change timeRecovery and scale operations stay within the operational window

Cloud cost follows these paths too. Keep compute, object storage, requests, data transfer, observability, and engineering operations as separate lines in the model. Cross-zone replication, cross-region movement, private endpoints, and egress can have different pricing rules; use the provider’s current pricing page for the regions and paths in scope instead of carrying a remembered rate into the forecast.

Compatibility is a capacity input because an unsupported client behavior creates operational work. Inventory client versions, producer acknowledgements, idempotence and transactions, compacted topics, consumer-group behavior, Kafka Streams, Connect, admin tools, authentication, quotas, and monitoring integrations. A design that saves storage but forces a rewrite of a high-volume producer has changed the project being priced.

Migration capacity belongs in the same worksheet. Record the topics that move first, the read and write paths during cutover, the offset or replay strategy, the rollback point, and the extra capacity needed while both systems are active. “The bucket is ready” is not a migration gate; the application contract and the recovery path must be ready too.

4How AutoMQ changes the operating model

Once the neutral framework is written down, a Kafka-compatible shared-storage platform becomes a concrete option to test. AutoMQ keeps the Kafka protocol and semantics at the client boundary while using a Shared Storage architecture for durable stream data. Its S3Stream storage library, WAL layer, data cache, and S3-compatible object storage form a path that can be measured with the same worksheet above.

The architecture changes what a broker-size decision means. Compute can be sized for request and partition work while retained bytes live in shared storage. Adding or replacing a broker does not require treating its local disk as the source of truth for every partition. That can reduce data-movement pressure during scaling, but the claim still needs a workload test: measure request capacity, cache behavior, WAL flushes, object-store traffic, and recovery time under the conditions that matter to your team.

WAL choice is part of the plan, not a footnote. AutoMQ documents S3 WAL, EBS WAL, Regional EBS WAL, and NFS WAL as storage options with different latency, infrastructure, and failure-domain implications. The WAL storage documentation should be read alongside the deployment model and the target workload. Record the selected backend, its zone behavior, its recovery path, and the measurements that justify its cache and bandwidth settings.

The object store is still a production dependency. Plan bucket lifecycle and retention separately from broker scaling, watch upload and fetch rates, and validate request limits and access policies in the selected cloud or S3-compatible service. Network locality matters as well: measure the path from clients to brokers and from brokers to storage, then verify the zones, endpoints, IAM permissions, and egress rules that make the path real.

This is the useful distinction between a product claim and an operating model. AutoMQ can provide the Kafka-compatible Shared Storage architecture; the platform team still owns the thresholds, dashboards, failure drills, and rollout gates. The evidence should tell you whether the architecture fits your workload before a larger migration makes the answer expensive to change.

5Decision checklist and FAQ

Use this checklist at a design review, before a pilot, and again before production traffic moves:

  1. Workload contract: ingress, egress, retention, replay, partition distribution, compression, and peak shape have owners and measurement windows.
  2. Byte-path model: broker, cache, WAL, object storage, metadata, and network limits are separate lines with units and thresholds.
  3. Failure evidence: broker replacement, object-store degradation, replay surge, zone loss, and metadata recovery have been tested or explicitly scheduled.
  4. Compatibility inventory: clients, transactions, compaction, Consumer groups, Connect, Streams, admin tools, authentication, quotas, and metrics are mapped.
  5. Cost model: compute, storage, requests, transfer, observability, and migration operations use current provider inputs and stated assumptions.
  6. Rollout gate: the pilot has a stop condition, rollback point, dashboard owner, and review date.

Diskless Kafka production readiness scorecard

5.1Does diskless Kafka mean the cluster has unlimited capacity?

No. Object storage can provide a different scaling boundary for retained bytes, but brokers, caches, WAL, object-store requests, metadata, and network paths still have limits. Capacity planning moves those limits into a visible worksheet.

5.2Is diskless Kafka the same as Kafka Tiered Storage?

No. Tiered Storage commonly keeps an active local tier and moves older segments to remote storage. A shared-storage design makes durable stream data independent of a broker’s local disk. Confirm the exact semantics and supported features of the implementation under evaluation.

5.3What should be measured before a pilot?

Measure the workload contract and the byte paths: producer and consumer rates, peak shape, retention, replay, cache behavior, write-buffer lag, object-store latency and requests, metadata recovery, network boundaries, and client compatibility. The measurements should match the failure and rollout gates you intend to use in production.

5.4How should AutoMQ appear in the capacity model?

Model AutoMQ as a Kafka-compatible Shared Storage option with an explicit WAL backend, cache policy, object-store configuration, network placement, and recovery procedure. Compare those inputs with the same workload and governance requirements used for the current Kafka design.

The original question was “How much capacity does a diskless Kafka cluster need?” The answer is a traceable set of byte paths, failure limits, and owners. If your current plan still makes broker count carry the weight of retained data, replay, and recovery, run this worksheet against one production-shaped workload. Then start an AutoMQ evaluation with the measurements and stop criteria in hand.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.