Table of Contents
Table of Contents
A diskless Kafka retention policy is not a longer value in retention.ms. It is a contract about which records must remain replayable, how quickly a consumer may read them, where the bytes live, and what happens when deletion or recovery falls behind. Once Kafka uses shared object storage as its durable layer, that contract becomes visible across storage, cache, metadata, network, and operations.
The common mistake is to copy a local-disk policy into an object-storage-backed design. A team keeps the same time window, assumes the same replay behavior, and discovers later that cold reads, object requests, cache eviction, or delayed deletion have changed the workload. Retention has become an architecture decision.
The Apache Kafka community's KIP-1150 discussion on diskless topics is useful context for that architectural shift. It is a proposal discussion, so treat it as design context and validate the behavior of the platform you operate.
The useful approach is to choose policy from the replay contract outward. Start with the work a retained record must support, translate that work into storage and read-path requirements, then test the failure and cost boundaries before increasing the window. The rest of the platform can follow from those choices.
1Retention starts with a replay contract
A retention window only has meaning when it is tied to a consumer action. “Keep the topic for a long time” does not tell an operator whether the oldest record must be available for an incident replay, a backfill, an audit request, or an offline export. Those uses have different latency, concurrency, and isolation requirements.
Write the contract in terms that an SRE and an application owner can both test:
| Workload | What must remain true | First measurement |
|---|---|---|
| Operational recovery | A consumer can restart from a known offset after a bad deployment or downstream outage | Time to resume and catch up from the selected offset |
| Data backfill | A replay job can read a historical range while head-of-log traffic continues | Cold-read throughput, consumer lag, and hot-path latency during replay |
| Audit or investigation | Records remain discoverable for the required period with controlled access | Retained-byte inventory, access logs, and deletion evidence |
| Archive or lakehouse export | Historical data can be read by the archive path without pretending it is an online Kafka replay | Export completion time and downstream storage state |
This classification prevents one topic default from serving incompatible jobs. An operational stream may need a narrow replay window with predictable reads. An audit stream may need a longer window, but its access pattern can tolerate a separate archive path. A backfill-heavy stream needs an explicit test for old-offset reads because its retention cost is partly a read cost.
The contract also needs an owner. Retention changes should identify the team that approves more history, the team that owns replay load, and the operator who verifies deletion. Without those names, a size cap or time cap becomes a suggestion rather than a control.
2What changes when durable data leaves the broker
In traditional Kafka, a broker owns local log segments and replicas. The retention worker deletes local segments after the configured policy allows it, and the broker's disk is the physical boundary that operators watch. A diskless design moves durable stream data to shared object storage. Brokers still handle Kafka requests, leadership, metadata, and serving caches, but a broker-local volume is no longer the source of truth for the retained history.
That shift changes the retention path in three connected ways:
- Data objects: Records are written into the shared storage path and organized into objects that can be read after a broker is replaced. Deleting a Kafka range must eventually remove or retire the corresponding durable objects.
- Serving state: A broker cache and any write-ahead log serve hot data and recent writes. Cache size affects latency and read isolation, but it is not the retention boundary.
- Metadata: Kafka offsets, segment or stream mappings, and deletion state tell readers which data is valid. A successful policy decision therefore needs metadata progress as well as object cleanup.
The distinction matters during a broker failure. If the retained history is tied to one broker's disk, recovery begins with disk attachment, replica recovery, or partition movement. If the history is in shared storage, another broker can serve the logical partition after it reconstructs the required metadata and warms the relevant cache. The recovery test still matters, but its variables have moved.
Retention deletion also deserves a separate observation. A topic can reach its logical time or size boundary before every remote object has been physically deleted. Object compaction, asynchronous cleanup, and storage lifecycle rules may create a gap between “record is no longer readable through Kafka” and “the last object byte is gone.” Treat those as separate states and measure both.
3Translate Kafka settings into policy boundaries
Kafka exposes retention controls such as retention.ms and retention.bytes, along with cleanup policy settings for delete and compacted topics. The configuration names remain familiar in a diskless Kafka deployment, but their consequences are distributed across the streaming storage layer and object store. The Apache Kafka topic configuration reference is the right starting point for the protocol-facing behavior.
Use a policy table that explains why each boundary exists:
| Boundary | Question to answer | What to verify in shared storage |
|---|---|---|
| Time | How far back must a consumer be able to read? | Offset availability, object age, and deletion eligibility |
| Size | What retained-byte limit protects the service or budget? | Logical bytes, compression, object overhead, and cleanup lag |
| Cleanup mode | Is history deleted, compacted, or both? | Compaction output, tombstone handling, and replay semantics |
| Read contract | How fast must an old offset be served? | Cache miss behavior, object reads, concurrency, and lag recovery |
| Access boundary | Who can replay or export the history? | ACLs, identity, storage permissions, and audit records |
A time policy and a size policy can point in different directions. A low-throughput topic may retain a long calendar window without reaching its byte cap. A high-throughput topic can hit its cap before its incident-recovery window expires. The effective policy is the boundary that arrives first, so operators should alert on both rather than watching only the time setting.
For planning, calculate logical retained bytes before adding architecture multipliers:
logical retained bytes
= ingest rate
× retention window
× compression effect
Then add the dimensions that the selected architecture actually charges or constrains: object storage, requests, cache, WAL, network, compaction, and any replicas or redundancy layers. The formula is a planning aid, not a promise of a fixed bill. Use the cloud provider's current pricing and the workload's measured compression and replay behavior.
Do not hide compaction inside a generic “retention” label. A compacted topic retains the latest value per key according to its cleanup rules, but its write, tombstone, and compaction workloads are different from an append-only topic. A policy review should name the cleanup mode and describe what a consumer is allowed to replay.
4Test the failure and cost boundaries
A policy is production-ready when it explains what operators will see while the system is under pressure. Shared storage removes some broker-local coupling, but it does not remove backlogs, object-store throttling, stale metadata, or runaway replay jobs. The tests should follow the contract rather than a generic benchmark.
| Scenario | Failure question | Evidence to keep |
|---|---|---|
| Broker replacement | Can another broker serve the retained range without copying the full history? | Recovery timeline, metadata convergence, and consumer lag |
| Object-store slowdown | What happens to cache misses and catch-up reads when storage responses slow down? | Request latency, retries, queue depth, and hot-read impact |
| Slow consumer | Does a lagging group evict hot data or create unbounded read pressure? | Cache hit ratio, eviction reasons, and producer or tailing-read latency |
| Deletion backlog | Can policy state advance while old objects remain? | Kafka-visible end of retention, cleanup backlog, and object inventory |
| Replay storm | Can multiple backfills share the cold path without starving production consumers? | Per-group read rate, object requests, and isolation controls |
The cost review should be equally explicit. Retained bytes are the visible term, but request charges, retrieval, network placement, compaction work, and cache capacity can dominate when replay becomes routine. A cost model should separate:
- bytes stored for the retention contract;
- bytes read during normal consumption and replay;
- object requests caused by the read and cleanup paths;
- broker compute and cache reserved for serving those reads;
- network traffic between clients, brokers, and storage;
- operator work for policy changes, audits, and recovery drills.
This is where a FinOps framework for long-retention Kafka storage growth can complement the topic-level policy. The accounting model does not replace a replay test, but it makes the owner of each cost driver visible.
Security belongs in the same review. A retained object is still business data after it leaves a broker. Check who can read historical offsets, who can trigger exports, which storage identity performs cleanup, and how deletion evidence is retained. A policy that meets a calendar requirement but cannot prove access control or deletion state is incomplete.
5Tiered storage and shared storage answer different questions
Kafka tiered storage and shared-storage Kafka are often discussed together because both use remote storage. They place the retention boundary in different parts of the system.
Tiered storage moves older log segments to a remote tier while the broker continues to own a local tier and the hot path. That can relieve local disk pressure and extend history. It does not automatically remove the broker's storage responsibilities, and remote reads still need a clear test for cache behavior, broker pressure, and recovery.
The boundary is easy to miss in an architecture review. A cloud-native Kafka checklist for remote log storage boundaries can help teams record which bytes remain broker-local, which reads cross the remote tier, and which recovery steps still move data.
Shared storage makes durable stream data a first-class shared layer. Brokers serve Kafka operations and cache data, while the storage layer holds the durable history. The operating question shifts from “Which broker owns this segment?” to “Which broker can serve this range, and what metadata and cache state does it need?” The answer can reduce data movement during replacement or scaling, but it also makes object layout, compaction, and remote-read observability part of the platform contract.
Use the comparison to frame an evaluation, not to declare a universal winner:
| Architecture | Retention boundary | Main test |
|---|---|---|
| Broker-local Kafka | Local disks and replica placement | Disk growth, reassignment, and replay pressure |
| Kafka tiered storage | Local hot tier plus remote historical tier | Remote fetch impact and tier coordination |
| Shared-storage Kafka | Shared durable objects plus serving cache | Metadata, object reads, cache isolation, and cleanup |
If a workload only needs a short online window, the additional shared-storage controls may not justify a change. If the workload combines long history, frequent replay, and elastic broker operations, the broker-local boundary deserves a direct comparison against shared storage. The decision follows the contract, not the label.
6How AutoMQ changes the operating model
Once the neutral policy is clear, AutoMQ fits as a Kafka-compatible shared-storage option. AutoMQ uses S3Stream, a shared streaming storage layer, so durable data can live in object storage while brokers handle Kafka protocol processing, scheduling, and cache-backed reads. The S3Stream architecture documentation describes the storage API and object-backed layout.
For retention, the important change is the boundary between durable history and serving capacity. A broker cache is sized for the access pattern, while the durable object layer holds the retained stream. A longer policy can therefore be evaluated as a storage and read-path decision instead of automatically becoming a larger broker volume. Broker replacement and scaling still require metadata and health checks, but they do not begin with copying every retained byte between local disks.
AutoMQ does not make retention free or erase operational choices. Teams still need to model object requests, cache sizing, compaction, cleanup lag, network placement, and access controls. The value of the architecture is that those variables are visible and separable, which gives an operator more than one lever when retention grows.
A practical evaluation should compare the same contract on both architectures:
- Produce a representative workload with the intended cleanup policy.
- Read from the head while a separate consumer replays the oldest retained range.
- Replace a broker and record when the replacement broker can serve the same offsets.
- Trigger policy cleanup and compare Kafka-visible expiry with object cleanup evidence.
- Attribute storage, requests, compute, cache, and network to the topic or workload owner.
Keep the measurements with the policy record. A retention value without its replay and cleanup evidence is a configuration choice, not a production design.
7A rollout checklist for retention policy
Start with a small set of policy classes rather than editing every topic default. Each class should name the replay purpose, time and size boundaries, cleanup mode, read expectation, owner, and deletion evidence. Topics that do not fit a class should stay in review until an owner writes the missing contract.
Before rollout, verify:
- Kafka clients and tools preserve the required offset, ACL, transaction, Connect, and Streams behavior.
- The oldest-offset replay test passes while tailing consumers and producers are active.
- Cache and object-store metrics show whether a replay is isolated from hot traffic.
- Cleanup state is observable from both Kafka and object storage.
- The cost model includes retained bytes, reads, requests, compute, cache, and network.
- Recovery drills cover broker replacement, storage slowdown, and an interrupted replay.
- Access and deletion evidence meets the team's security and governance requirements.
After rollout, treat retention as a managed resource. Review policy changes with the workload owner, record why the window changed, and compare measured replay demand with the original contract. When a topic stops needing online replay, move it to an archive workflow rather than letting a default window grow without review.
8FAQ
8.1Does diskless Kafka make retention unlimited?
No. Object storage can provide elastic capacity, but a retention policy still creates storage bytes, read work, requests, metadata, and governance obligations. The practical limit comes from the replay contract and the cost and recovery behavior you can operate.
8.2Should every topic use both time and size limits?
Not automatically. A time limit protects the replay window, while a size limit protects capacity or budget. Use both when the workload needs a calendar guarantee and a hard byte boundary, then alert when the effective policy is the earlier boundary.
8.3Is tiered storage the same as diskless Kafka?
No. Tiered storage adds a remote historical tier to a broker-centered design. Diskless Kafka places durable stream data in shared storage and treats brokers as serving and coordination nodes. Compare the read path, recovery path, and cleanup evidence for your workload.
8.4How do I choose a retention window?
Start with the replay jobs that must succeed, then record their oldest offset, acceptable catch-up time, concurrency, and access controls. Convert that contract into time, bytes, cleanup, and cost boundaries, and validate it with a cold-read test.
A retention policy is doing its job when an operator can answer four questions without guessing: which records must remain, who will replay them, where the bytes are, and how deletion is proved. That is the difference between a number in a topic configuration and a production contract for diskless Kafka.
If you are evaluating that contract on a Kafka-compatible shared-storage platform, run the same replay, failure, and cleanup checks against your own workload. You can try the AutoMQ open-source code and use the results to decide whether the retention boundary should remain attached to brokers or move into a shared storage design.
