Blog

Diskless Kafka Kafka Streams: A Production Framework

Table of Contents

Table of Contents

A Kafka Streams application can keep producing correct results while its Kafka cluster is changing storage architecture underneath it. That is exactly why a diskless Kafka evaluation needs more than a broker benchmark. The application has state stores, changelog topics, offsets, repartition topics, and a restoration path, and each one can fail or slow down in a different place.

The first question for a platform team is not whether Kafka Streams can connect to an object-storage-backed cluster. It is whether the team can explain where each byte becomes durable, which state is rebuilt after a task moves, and what recovery looks like when a broker, an application instance, or an object-storage path is unavailable. A useful production framework keeps those questions separate, then measures the handoffs between them.

Kafka Streams decision map for Diskless Kafka

1What Kafka Streams means in a Diskless Kafka design

Kafka Streams is a client library that turns Kafka topics into a processing topology. A topology may keep a key-value state store for joins, aggregations, windows, or deduplication; it may also create repartition topics so that records arrive at the task that owns the relevant key. The state store belongs to the Kafka Streams application instance. The input, repartition, and changelog topics belong to Kafka. A diskless design changes the storage path for those topics, but it does not automatically move application state into object storage.

That distinction prevents a common architecture error. Replacing broker-local log storage with shared storage can make retained Kafka records independent of a particular broker, while a Streams task still needs a local state store for its active reads and writes. The state store is usually restored from its changelog topic when a task starts or moves. The changelog is therefore the durable bridge between application state and the Kafka cluster, even when the broker stores that topic through an object-storage-backed path.

A production review should trace the contracts separately:

  • Kafka log contract: records, offsets, ordering, retention, compaction, and producer acknowledgment for input, repartition, and changelog topics.
  • Streams state contract: the state store format, cache and flush policy, task ownership, standby strategy, and restore behavior.
  • Recovery contract: the time and network path required to rebuild a task, catch up its changelog, and resume output with the expected semantics.

The Kafka Streams developer guide describes the state-store and topology concepts that operators should map to their own applications. The guide is a starting point, not a production proof. A state store that looks small in a design document can become the dominant recovery cost when a window, join, or repartition step retains a large working set.

2Trace the input, state, and restoration paths

The easiest way to make Kafka Streams measurable is to draw the paths that a record can take. A source record enters an input topic and is read by a Streams task. The task may update a local state store, emit a record to a downstream topic, or write to a repartition topic before another task processes it. The state store can be backed by a changelog topic, and the output topic can be consumed by another application or connector. Those are different streams with different retention and recovery requirements.

PathWhat movesWhat to measureFailure question
Input and repartitionRecords from source topics into tasksFetch latency, consumer lag, partition distribution, and network bytesCan tasks keep up when a partition or broker moves?
State storeLocal updates, cache flushes, and persistent store filesStore size, cache hit ratio, flush time, disk or memory pressure, and task throughputCan the instance keep serving while the store is restored or compacted?
ChangelogState updates written to Kafka for recoveryProduce latency, retention, compaction behavior, and restore bytesIs the changelog complete and readable after task loss?
OutputProcessed records and transactional boundariesEnd-to-end latency, commit rate, retries, and downstream lagDo retries preserve the required processing semantics?

This table is intentionally more detailed than a broker throughput chart. Kafka Streams performance is shaped by the largest state store, the hottest key distribution, and the slowest restoration path, not only by source-topic ingress. A diskless Kafka cluster can make broker storage elastic while leaving the application instance short of local cache or recovery bandwidth. Both sides have to be measured before a capacity decision is made.

The state store and the changelog also have different lifecycles. A local store can be deleted and rebuilt from the changelog, but that rebuild consumes broker fetch capacity, object-storage reads, network bandwidth, and application CPU. A changelog topic configured for compaction may retain the latest record for each key, but compaction timing and restore behavior still need to be tested under the Kafka release and topology settings in use. Retention, compaction, and task placement are operational choices, not properties that appear automatically when a topic is called a changelog.

Diskless Kafka data path for Kafka Streams

3The mechanism: brokers, cache, metadata, and object storage

A shared-storage Kafka architecture changes the broker side of this picture. The broker continues to handle Kafka protocol requests, partition metadata, access control, and fetch scheduling. The durable topic bytes can move through WAL (Write-Ahead Log) and into S3-compatible object storage, while caches serve the hot path. The Streams application still owns task execution and state-store behavior, so the architecture has two caches and two recovery conversations: broker-side data access and application-side state restoration.

That is why “Kafka on S3” is not a sufficient description. An object store can hold durable topic data, but Kafka Streams still depends on predictable fetch latency, correct offsets, and a changelog that can be consumed in the order the state store expects. Request rate, object layout, compaction, network locality, and the selected WAL storage affect the broker path. Store format, cache sizing, and task assignment affect the application path.

The difference from Kafka Tiered Storage matters here. KIP-405 describes a remote tier for older log data while an active local tier remains part of the broker storage model. KIP-1150 describes a diskless-topic direction in which object storage can replace durable broker storage while local disks serve as cache. Both references are useful for framing the design, but the release, implementation, and topic capabilities in your target environment still need verification.

The practical implication is simple: a diskless topic can change the cost and recovery behavior of a Streams input or changelog, but it does not make the Streams state store remote by definition. If a task loses its local store, it still has to restore from Kafka. The review should therefore record the source of truth for each layer:

LayerDurable sourceWorking setRecovery owner
Kafka topic dataWAL and object storage, according to the implementationBroker cache and request buffersKafka cluster and storage operators
Streams state storeChangelog topic plus application store filesLocal disk or memory on the Streams instanceStreams application owner
Task metadataKafka consumer-group and topology metadataTask assignment and runtime stateKafka Streams application and cluster metadata layer
Output and side effectsOutput topic, transaction log, or external sinkProducer and sink buffersApplication and sink operators

A plan that puts every row under “object storage” has collapsed the architecture too early. The rows are connected, but their failure signals and owners are different.

4Failure, cost, and compatibility checks

Start the failure review with task movement rather than broker failure alone. A broker can disappear while a Streams task remains healthy, or a Streams instance can restart while all brokers remain available. The platform needs evidence for both cases, plus the combined case where a task is restoring while the cluster is serving normal traffic.

ScenarioEvidence to collectProduction gate
Broker replacement during processingInput fetch latency, changelog fetch latency, object-store requests, and task lagBroker recovery does not push the topology beyond its lag or latency objective
Streams instance restartState-store restore bytes, restore duration, CPU, local storage, and output behaviorThe task resumes within the recovery window and produces the expected records
Changelog compaction or retention pressureCompaction lag, tombstones, retained bytes, and restore correctnessA new task can rebuild the required state from the available changelog
Hot-key or skewed partitionRecords per key, task assignment, store growth, and repartition trafficThe busiest task has headroom without starving other tasks
Object-storage or endpoint degradationFetch and produce retries, request errors, upload lag, and network pathBackpressure and alerts expose the degraded dependency before data loss
Rolling upgrade or topology changeRepartition behavior, state-directory compatibility, and offset continuityThe topology change has a tested rollback and no silent state reset

Cost follows the same boundaries. Keep broker compute, WAL or cache resources, object-storage capacity, object-storage requests, application-instance storage, network transfer, and operations as separate lines. A changelog may be compacted, but it still creates write and read requests; a state store may be local, but it still consumes disk, memory, and recovery bandwidth. Use the provider’s current pricing for the selected region and endpoint paths instead of carrying a generic “S3 has a lower unit cost” assumption into a business case.

Compatibility is a separate gate from storage economics. Inventory the Kafka client and Streams versions, producer acknowledgment and transaction settings, state-store types, log compaction, repartition topics, standby replicas, interactive queries, security configuration, metrics, and deployment model. The topic API may look familiar while an application depends on a corner of task restoration or transaction behavior that the target implementation has not verified.

A useful pilot records the same workload in two views. The Kafka view follows bytes and offsets through input, repartition, changelog, and output topics. The Streams view follows tasks, state-store size, restore progress, cache pressure, and output commits. If those views disagree about where the backlog lives, the pilot is not ready to support a production decision.

5How AutoMQ changes the operating model

Once the neutral framework is explicit, AutoMQ fits into a concrete category: a Kafka-compatible streaming platform that replaces broker-local durable log storage with a Shared Storage architecture. AutoMQ documentation describes reuse of the Apache Kafka compute layer and S3Stream as the storage layer, with WAL storage providing a low-latency write path before data reaches object storage. The existing Diskless Kafka architecture guide explains which responsibilities remain in the broker, while the Kafka compatibility guide is useful when turning a topology inventory into a client test plan.

That cut point is relevant to Kafka Streams because the application can continue to use Kafka topics, offsets, consumer groups, and the Streams client model while the broker’s durable topic path changes. Input, repartition, changelog, and output topics can be evaluated through the same compatibility and recovery worksheet. The application’s state store remains an application concern, so the pilot still needs to size local state directories, cache, restore bandwidth, and standby policy.

The WAL choice belongs in that worksheet. AutoMQ’s WAL storage documentation describes WAL as the low-latency persistence and failover-recovery layer, with storage options that vary by deployment and latency requirement. Record the selected WAL type, its failure domain, object-storage endpoint, cache policy, and recovery procedure. Do not infer a Streams state-store latency target from the broker WAL description; they are different paths.

The compatibility claim still needs workload proof. AutoMQ’s documentation reports Kafka protocol and semantic compatibility and references Kafka Streams and system testing. A platform team should turn that claim into its own test matrix: run the exact topology, include state restoration, exercise compaction and repartition topics, test transactional boundaries if used, and compare lag and output correctness during broker and application events. This preserves the value of Kafka compatibility while keeping the burden of proof at the application boundary.

Kafka Streams production readiness scorecard

A migration can then proceed in gates. First, mirror or replay a representative input window and compare offsets, state-store materialization, and output records. Next, restart a Streams instance and replace a broker while measuring restore and fetch paths. Finally, exercise the rollback point with the same security, schema, and transaction settings used in production. Shared storage can reduce broker data movement, but it does not remove the need to rehearse state restoration.

6Decision checklist and FAQ

Use this checklist when a platform team reviews Diskless Kafka for a Kafka Streams workload:

  1. Separate the sources of truth. Document which bytes are in Kafka topics, which state is in local stores, and which metadata controls task ownership and offsets.
  2. Measure restoration. Record state-store size, changelog fetch rate, restore duration, local disk or memory pressure, and output behavior during restart.
  3. Test the difficult topics. Include repartition topics, compacted changelogs, transactional output, skewed keys, and long replay windows where the topology uses them.
  4. Model the full cost path. Keep broker compute, WAL, object storage, requests, application storage, network, and operations separate.
  5. Define failure gates. Set explicit limits for broker replacement, Streams restart, object-storage degradation, compaction lag, and rollback.
  6. Validate compatibility in context. Use the real topology and client versions, then record which Kafka and Streams behaviors were tested and which remain assumptions.

6.1Does Diskless Kafka remove Kafka Streams state-store storage?

No. Diskless Kafka changes how Kafka topic data is durably stored at the broker layer. A Kafka Streams task can still use local state stores, caches, and state-directory files. The changelog topic provides the recovery path, so the restore test remains a first-class production gate.

6.2Are changelog topics safe on object storage?

They can be, when the Kafka implementation preserves the topic’s ordering, offset, compaction, retention, and recovery semantics. Verify the supported topic behavior in the release you plan to run, then test a full state-store restore after broker and application failures.

6.3Is Kafka Tiered Storage the same as Diskless Kafka?

No. Tiered Storage normally moves older log data to a remote tier while the active broker-local tier remains. A diskless or Shared Storage design changes the durable ownership boundary for active topic data. Kafka Streams still needs its own state-store and changelog checks in either architecture.

6.4Which metrics matter first?

Start with input and changelog fetch latency, task and consumer lag, state-store size, restore duration, cache and local-storage pressure, repartition traffic, object-storage request errors, and output commit behavior. Add cost and network metrics that match the endpoints and availability zones in the deployment.

The opening question was whether Kafka Streams can remain predictable when Kafka’s durable log moves away from broker-local disks. The answer is measurable: keep the application state contract visible, test the changelog restore path, and treat broker storage and task storage as separate failure domains. If your current design passes those gates, start an AutoMQ evaluation with the topology, recovery window, and stop criteria already written down.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.