Table of Contents
Table of Contents
Kafka retention is straightforward to describe and hard to review. A topic can keep records for a period, but the setting does not tell you whether a replay will hit a local SSD, a remote tier, a cache, or an untested recovery path. On Google Cloud, that distinction affects the data you recover, backfill latency, and billing lines for old reads.
The useful question is not “how long should we retain this topic?” It is “which replay promise are we making, and which storage path can prove it?” This guide treats GCP Kafka retention as a contract between data age, replay frequency, recovery evidence, and cost. It gives architects and FinOps teams a way to compare local SSD, managed remote storage, and a Kafka-compatible Shared Storage architecture without assuming that one tier is right for every topic.
1Retention starts with a replay contract
Retention is often chosen from a calendar: a short window for application events, a longer window for analytics, or an extended window for audit. Calendars are useful only after the team names what a reader must do with the retained data. A topic used for continuous catch-up has a different storage shape from one replayed once after a schema migration.
Write the contract before comparing storage options. It should state:
- Replay horizon: the oldest record a supported job must read, expressed as an offset or event time rather than a vague “long term.”
- Replay objective: whether the consumer must catch up at tailing speed, meet a backfill freshness target, or complete without data gaps.
- Recovery boundary: what must remain readable after a broker, zone, network path, or storage dependency fails.
- Evidence: the producer ledger, offsets, storage events, and consumer output that prove the contract in a test.
This contract changes the sizing conversation. Retention bytes describe how much data exists, while replay behavior determines how much is hot, how often it is fetched, and which path must be provisioned. A replay job still needs read bandwidth, permissions, and schema context to make remote history useful.
2Local SSD is a latency choice with a capacity boundary
Local SSD keeps recent Kafka data close to the broker. That can make tailing reads and short catch-ups straightforward because the active log and its cache share a fast local path. The trade-off is structural: the retained working set is bounded by the local capacity attached to the broker layout, and growth can turn a retention change into a capacity or rebalance event.
For a GCP Kafka deployment using local SSD, review the complete loop rather than the device label:
- Measure the hot window that consumers actually read, including bursts after a downstream outage.
- Reserve headroom for compaction, replication or recovery work, and the largest replay that must run while normal traffic continues.
- Test what happens when a broker or zone is replaced and the local copy is not immediately available.
- Reconcile local storage, compute, and inter-zone traffic lines with the retention policy.
Local SSD is a sensible fit when the replay contract is short, frequent, and latency-sensitive. It becomes a poor default when teams retain data speculatively but rarely read it. Paying to keep every retained byte in the hot path can hide the real cost driver: the size of the history, not the speed of the last few hours.
3Remote storage separates history from the broker, but replay still has a path
Remote storage changes the capacity boundary. Older segments can live outside the broker’s local working set, letting the topic retain more history without treating every byte as hot broker capacity. Google Cloud’s Managed Service for Apache Kafka pricing separates compute, local SSD, remote storage, and network dimensions. That separation is useful for a cost model, but it does not answer whether a replay meets your application objective.
Ask three questions about the remote path:
- Where does the read start? A current consumer may read from a cache, while an old offset triggers a remote fetch and a different latency profile.
- What happens during a dependency failure? A denied bucket, endpoint outage, or throttled request can turn a healthy retention setting into a replay incident.
- Who owns the recovery evidence? The platform team needs request logs, offsets, error behavior, and a replay result that can be compared with the producer record set.
Remote storage is therefore a replay design, not a checkbox. A test that reads only the newest records proves the hot path. It says very little about a backfill that starts weeks earlier, crosses a compaction boundary, or runs after a broker replacement. The test should include a cold range and the identity that a recovery environment will use.
4Compare options with the same workload questions
Each option should answer the same questions. The matrix is qualitative; a production estimate must use the current configuration and price sheet.
| Review question | Local SSD | Managed remote storage | Shared Storage architecture |
|---|---|---|---|
| What stays close to the broker? | Recent log data and cache | A configured hot window and cache | WAL and cache for active traffic; durable stream in object storage |
| What makes long retention expensive? | Broker capacity and replacement headroom | Remote storage, requests, and cold reads | Object-storage bytes, requests, WAL choice, and network path |
| What must a replay test include? | Local capacity during catch-up and broker replacement | Remote fetch, identity, quotas, and historical offsets | WAL recovery, object-storage reads, metadata ownership, and cache refill |
| When is the option a good fit? | Frequent short replays with tight latency goals | Large history with a defined remote-read contract | Independent compute and durable storage with an explicit recovery test |
The table does not rank the options. It exposes where the operational work moves. Local SSD concentrates capacity and recovery pressure on brokers. Remote storage moves more of that pressure into service boundaries and cold-read behavior. Shared Storage architecture separates broker compute from the durable stream, but it still requires a named WAL type, object-storage access, metadata fencing, and replay evidence.
5Make replay a test, not a promise
Replay is the moment when retention policy meets reality. Pick a representative Topic and create a ledger of record keys, timestamps, offsets, and producer outcomes. Then run a fresh Consumer from an older offset while normal traffic continues. Keep the output so another operator can check for gaps, duplicates, ordering changes, and application-level failures.
At minimum, test these paths:
- Cold read: start beyond the hot window and record fetch latency, remote requests, and cache refill behavior.
- Broker replacement: isolate or replace the broker that owns the tested Partition, then repeat the same offset range.
- Storage error: deny or interrupt the remote path in a controlled environment and verify that the Consumer receives an explicit failure rather than silently skipping data.
- Backfill pressure: run the replay while producers and tailing Consumers maintain their normal workload, then compare lag and recovery signals with a no-fault control.
The pass line should be written in terms of the contract. “The replay completed” is incomplete if it dropped records or required manual repair. Name the tested range, observed gaps or duplicates, workload time budget, and evidence retained for storage and metadata paths.
6Cost follows bytes, reads, and failure behavior
Retention cost is not one number. A useful worksheet separates retained bytes, storage requests, compute, inter-zone traffic, egress, and recovery reads. The same topic can have a low average read rate and a costly replay month if a backfill scans a large cold range or retries after an endpoint error.
Use counters instead of generic savings language. For each workload, record:
- bytes retained by age band, with hot and cold ranges labeled;
- bytes read during normal consumption and replay;
- request counts and retry counts for remote reads and writes;
- cross-zone or egress bytes created by the chosen topology; and
- the compute capacity held for peak replay rather than average traffic.
Google’s pricing page provides billing dimensions for its managed Kafka service, but it cannot know your replay frequency or failure behavior. Apply current regional rates to your workload counters. If the team has not run a cold-read test, label the estimate as an assumption.
7Where AutoMQ fits in the decision
The framework points to a broader architectural option when broker-local capacity is the wrong place to anchor long retention. AutoMQ is a Kafka-compatible cloud-native streaming platform built on a Shared Storage architecture. Its S3Stream storage layer keeps the durable stream in S3 storage, while WAL storage and Data caching serve the write and read paths according to the selected deployment.
That separation changes what the retention review must inspect. A broker replacement does not imply that the full retained history must be copied from one broker-local disk to another, but the replacement still needs metadata ownership, access to the selected WAL and S3 storage, and a cache path for tailing and catch-up reads. The architecture overview and WAL storage guide describe the layers to verify.
AutoMQ does not remove the replay contract. It gives the team a different set of boundaries to test: Kafka protocol behavior at the client, WAL durability, object-storage reads, KRaft metadata ownership, and Data caching. For a customer-owned cloud deployment, confirm the supported GCP storage, network, IAM, and deployment choices in the product documentation before sizing a cluster. A Shared Storage design is useful only when its recovery evidence is stronger than the local-disk assumption it replaces.
For related cost framing, see Kafka on GCP: Compute Engine, Managed Kafka, Pub/Sub, and Diskless Kafka and Retention Cost Breakpoints for Cloud-Native Kafka Estates.
8FAQ
8.1Is remote storage always the right answer for long retention?
No. Long retention still needs a replay objective, a cold-read test, and an owner for the remote path. A short, frequently replayed workload may prefer a larger hot window, while an audit stream may value capacity and recoverability over tail latency.
8.2Does a larger retention.ms guarantee replay?
No. retention.ms expresses when Kafka may remove records. It does not prove that the storage path, credentials, schema context, consumer offsets, or application state will be available when a replay starts.
8.3Is Shared Storage the same as Kafka Tiered Storage?
No. Kafka Tiered Storage can offload older log segments while the broker-local replica model remains. A Shared Storage architecture makes the durable stream a first-class shared layer and therefore tests WAL, object storage, metadata, and cache recovery together.
8.4What should FinOps request from the platform team?
Request the retention contract, age-banded byte counts, normal and replay read counters, request and retry counts, network bytes, and the latest cold-read test. Those artifacts are more useful than a single monthly total because they explain which workload behavior produced the spend.
When a topic is retained without a replay test, the storage choice is still an assumption. Start with one offset range, one cold-read run, and one failure drill. Then choose the path that can satisfy the contract and show its work. If you want to apply the same worksheet to a Kafka-compatible Shared Storage deployment, start an AutoMQ evaluation with the replay ledger and cost counters ready.
