Blog

Diskless Kafka KRaft Operations: A Production Framework

Table of Contents

Table of Contents

A KRaft controller can be healthy while an Apache Kafka cluster is still difficult to operate. The metadata quorum may elect a successor quickly, yet a broker replacement can remain blocked by replica movement, local disk capacity, or a long catch-up window. Diskless Kafka changes that boundary: KRaft still owns cluster metadata, while a shared storage layer owns durable topic data.

That separation gives platform teams a more useful question than “does KRaft work?” The question is whether controller decisions, broker recovery, and storage recovery can be measured independently. The answer determines how you design failure drills, capacity plans, and migration gates.

1KRaft is a metadata operating model, not a storage model

KRaft replaces Kafka’s external ZooKeeper dependency with a metadata quorum managed by Kafka controllers. The active controller writes metadata changes to the quorum, and the other controllers replicate those changes as followers. Apache Kafka’s KRaft documentation and KIP-631 describe the controller quorum, broker registration, and the relationship between controllers and brokers.

The distinction matters because KRaft does not decide where a topic’s records live. In a conventional Kafka cluster, the broker still owns local log segments and replicates partition data to other brokers. A controller operation such as reassignment therefore has two parts: commit the metadata change, then move enough bytes for the target replica set to become useful. A metadata leader election can finish while the data movement is still consuming disk, network, and operator attention.

Diskless Kafka separates those concerns. The controller quorum remains a metadata system; the topic log uses a storage layer that can place durable records in object storage, with a write-ahead path and cache for the latency-sensitive work. Apache Kafka’s KIP-1150: Diskless Topics frames diskless topics as a design in which object storage can serve as the primary durable layer rather than an archive for older segments.

The controller remains important. Its responsibility is easier to reason about when metadata availability and data availability are kept distinct, because they have different failure signals and recovery actions. Teams evaluating the underlying storage proposal can use this Kafka KIP review workflow to keep protocol semantics, recovery, and cost questions in the same decision record.

2The four surfaces an operations review must keep separate

The fastest way to create a confusing KRaft runbook is to treat every recovery symptom as a controller problem. A diskless Kafka review should keep four surfaces visible:

  • Metadata quorum: controller leadership, quorum replication, broker registration, topic and partition metadata, and the ability to commit a metadata change.
  • Broker compute: request handling, partition leadership, scheduling, cache management, and the process lifecycle of each broker.
  • Durable stream storage: the write-ahead log, object-storage objects, retention, and the consistency rules that make a record readable after a broker is replaced.
  • Network and identity boundaries: controller listeners, broker-to-storage paths, object-storage permissions, private endpoints, and the identities allowed to change cluster state.

The surfaces interact, but they should not share one undifferentiated alert. A controller quorum that cannot commit metadata is a control-plane incident. A broker with an empty cache after restart is a data-path warm-up event. An object-storage permission failure is a storage access incident. Each can affect producers and consumers, yet the first diagnostic question is different.

The separation also improves incident ownership. The platform team can ask whether metadata commits are progressing; the storage owner can ask whether durable objects and the write-ahead path are available; the network team can verify that the expected identity can reach the storage endpoint. A runbook that names the surface before naming the fix is easier to execute under pressure.

3What changes when a controller fails

In KRaft, one controller is the active leader and the other controllers act as hot standbys for the metadata quorum. When the active controller fails, a follower can take leadership if the quorum still has the majority needed to make progress. Brokers then reconnect to the current quorum leader and continue from the metadata log they have not yet applied.

The recovery target is therefore not “restart the controller.” It is “restore a quorum that can commit metadata and confirm that brokers have rejoined the current metadata view.” The operator should record the controller transition time, the period during which metadata commits were unavailable, and the time each broker took to re-register. These are control-plane measurements; they should not be blended with the time required to warm a cache or re-read records from object storage.

If the controller quorum loses its majority, metadata writes cannot safely commit. Existing data-plane traffic may continue for a period depending on the operation and the state already known by brokers, but changes that require metadata commits are constrained until quorum progress returns. A runbook should make that boundary explicit so that an incident commander does not mistake “some reads still work” for “the cluster is healthy.”

Controller listeners also deserve a separate check. KIP-631 recommends isolating controller ports from clients; clients should reach brokers, while administrative requests are forwarded to the controller quorum as needed. In a diskless deployment, this boundary becomes even more valuable because the storage endpoint and controller endpoint have different permissions and different failure domains.

4What changes when a broker fails

A failed broker still removes compute capacity and may remove partition leadership. The difference is what the platform must rebuild. With broker-local persistent logs, replacement usually means restoring a broker and copying partition replicas until the new broker can carry its durability role. With shared primary storage, replacement can focus on bringing up compute, restoring ownership and routing, and proving that the broker can read the durable stream through the storage layer.

A replacement is not instant or free. A cold broker has to establish credentials, reach the storage endpoint, load metadata, and warm the cache for its workload. A workload with replay-heavy consumers can create a very different recovery profile from a workload that reads only the newest records. The right measurement is not a generic recovery promise; it is the time from broker admission to the point at which the selected workload meets its own latency and throughput objectives.

The operational difference is visible in the failure matrix below. It describes the decision boundary rather than promising one universal recovery time.

EventFirst signalWhat the operator provesData movement question
Active controller lostController leadership change or metadata commit errorsA follower leads, the quorum can commit, and brokers re-registerNo topic-log copy is implied by the controller change
Controller quorum loses majorityMetadata commit stallsQuorum membership and network paths are restoredTopic data may remain readable, but metadata changes wait
Broker process lostBroker heartbeat or registration failureReplacement compute joins with the expected identity and metadataCache warm-up and read recovery must be measured
Storage endpoint unavailableWrite, fetch, or object access errorsThe write-ahead path, object store, and permissions recoverRecords may be durable elsewhere; verify visibility before retrying
Network policy changedConnection or authorization failuresController, broker, and storage paths are allowed separatelyDo not infer data loss from a blocked path

The table is useful during a drill because it prevents the team from applying a controller fix to a storage incident. It also makes a gap visible: if the runbook has no check for cache warm-up or object-store access after broker admission, it has not tested the diskless part of the design.

5A practical measurement loop for KRaft operations

Treat operations as a sequence of observable gates. The following loop works for a failure drill, a planned broker change, or a migration rehearsal:

  1. Establish the baseline. Capture controller quorum health, metadata commit behavior, broker registration, producer acknowledgement behavior, consumer fetch behavior, storage access, and network errors before changing anything.
  2. Inject one failure. Stop one controller, isolate one broker, deny one storage path, or remove one permission. Keep the scope narrow enough that the resulting signal has one plausible cause.
  3. Measure the control plane. Record when the controller quorum changes leadership, when metadata commits resume, and when brokers observe the current metadata view.
  4. Measure the data plane. Record producer and consumer symptoms, cache misses, write-ahead errors, object-store operations, and the point at which the workload returns to its baseline behavior.
  5. Check the boundary. Confirm that the observed recovery did not rely on an unrecorded manual change, a bypassed permission, or a silent data copy.
  6. Write the rollback condition. Define which signal stops the exercise and who owns the decision to restore the previous state.

The sequence forces the team to distinguish a fast metadata election from a complete service recovery. It also yields evidence that can be compared across architectures without using a vendor-specific score as a substitute for a workload result.

6Tiered Storage and Diskless Kafka answer different questions

Tiered Storage is often introduced as a synonym for diskless Kafka because both use object storage. The operating model is different. Tiered Storage keeps an active log on broker-local storage and moves older segments to a remote tier. It can reduce the amount of local retention a broker must carry, while broker-local state remains part of the write and recovery path.

Diskless Kafka uses a storage layer designed for shared primary durability. The broker still needs a latency-sensitive write path and cache, but the durable stream is not tied to one broker’s local disk. That changes the work required for broker replacement, scaling, and long-retention planning. It does not remove the need to validate object-storage latency, permissions, request behavior, or network reachability.

A design review should capture that difference in a decision table, not in a marketing label:

Review questionTiered StorageDiskless Kafka
Where is the active log written?Broker-local storage, with remote segments for older dataA shared storage layer with a write-ahead path and object storage
What does broker replacement restore?Local log state plus remote-tier awarenessCompute, metadata ownership, storage access, and cache readiness
What drives long-retention capacity?Local retention plus remote retentionObject storage retention plus write-ahead and cache sizing
What must a failure drill prove?Local disk health, segment offload, and remote readsQuorum progress, storage access, durable visibility, and cache recovery

There is no rule that every Kafka cluster should move to diskless storage. A stable cluster with modest retention and predictable growth may keep a local-storage design with a simpler runbook. The point is to choose deliberately: if controller operations repeatedly trigger data placement work, evaluate whether the storage architecture is the constraint.

7Where a Kafka-compatible shared-storage platform fits

Once the team has separated metadata, compute, storage, and network checks, it can evaluate implementations against the same gates. AutoMQ fits this category as a Kafka-compatible cloud-native streaming platform with Shared Storage architecture and stateless brokers. Its S3Stream storage layer uses WAL (Write-Ahead Log) and object storage for durable stream data, while brokers handle protocol requests, leadership, scheduling, and cache behavior.

That architecture does not change KRaft’s quorum rules. It changes what sits behind a controller decision. Adding compute capacity does not require treating every durable byte as a broker-owned asset, and replacing a broker can focus on re-establishing compute and storage access instead of rebuilding the entire durable log locally. The benefit should be demonstrated with the measurement loop above: controller convergence, broker admission, storage visibility, and workload recovery are separate observations.

The governance boundary is part of the same evaluation. In a customer-controlled deployment, AutoMQ BYOC documentation describes how the deployment runs in the customer’s cloud environment. The review should still verify the actual identity, endpoint, encryption, audit, and network controls used by the chosen object store. Shared storage changes the operating model; it does not waive the operator’s responsibility for access and failure boundaries.

8A production readiness checklist

Use the following checks before calling a diskless Kafka KRaft design ready for production:

  • Quorum: Document controller roles, quorum membership, listener isolation, reconfiguration procedure, and the signal that indicates metadata commits have resumed.
  • Broker lifecycle: Test registration, planned removal, unplanned restart, cache cold start, and the point at which the workload is considered recovered.
  • Storage path: Verify write-ahead durability, object-storage access, retention behavior, request errors, and recovery when the storage endpoint is unavailable.
  • Compatibility: Run representative producers, consumers, admin tools, transactions, consumer groups, connectors, and offset checks against the intended Kafka-compatible surface.
  • Observability: Keep controller, broker, cache, write-ahead, object-storage, network, and application symptoms queryable in one incident timeline.
  • Security: Test the identities that can alter controller state, operate brokers, and read or write durable stream data. Rotate credentials in a rehearsal, not during the first incident.
  • Rollback: Define the cutover owner, stop conditions, offset validation, and the path back to the prior architecture before migration traffic is enabled.

The checklist is intentionally operational. A slide that says “KRaft is enabled” says little about how the cluster behaves when a controller disappears, a broker returns cold, or the storage endpoint rejects a request. Those behaviors are what the production decision must accept. For a broader pre-production gate, compare these checks with the Diskless Kafka readiness checklist, then record which items are measured in your own workload.

9FAQ

9.1Does KRaft make Kafka diskless?

No. KRaft changes Kafka’s metadata quorum and removes the external ZooKeeper dependency. A diskless design changes the durable topic-data layer separately.

9.2Does a controller election recover a failed broker’s data?

No. The election restores metadata leadership. Whether a failed broker’s workload recovers depends on broker admission, storage access, cache state, and the behavior of producers and consumers.

9.3Is object storage enough to make a diskless Kafka design reliable?

No. Reliability depends on the full path: metadata quorum, write-ahead durability, object-storage consistency and permissions, cache behavior, network reachability, and an exercised recovery procedure.

9.4What should teams measure first?

Measure metadata commit recovery and broker re-registration separately from storage visibility and workload recovery. The separation shows whether an incident is in the controller, compute, storage, or network surface.

KRaft gives the cluster a clear metadata control plane. Diskless Kafka gives the platform a chance to make durable data independent from broker compute. The production test is the boundary between them: inject one failure, record what recovers first, and keep the rollback condition visible.

If your current runbook still treats every broker change as a data-copy project, run the same drill against a Kafka-compatible shared-storage design. The migration guardrails for diskless Kafka can help structure cutover and rollback questions, and the AutoMQ GitHub project is a practical starting point for testing the controller, broker, and storage checks in one environment.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.