Table of Contents
Table of Contents
A Kafka platform purchase is often reduced to a capacity worksheet: broker types, disk sizes, and a storage estimate. That worksheet is incomplete when the candidate is a diskless Kafka design. The important questions move underneath the product label: where an acknowledged record becomes durable, what brokers keep in local cache, how metadata controls ownership, and what happens when storage or network paths fail.
Apache Kafka®'s accepted KIP-1150: Diskless Topics captures the direction clearly. A diskless topic can use object storage for durable user data while broker disks serve as cache or short-lived working space. That does not mean that a production cluster has no disks. It means the procurement decision should measure the responsibilities of each layer instead of treating a diskless architecture as a literal absence of local media.
The useful deliverable is therefore a checklist with evidence gates. The team should be able to show which Kafka behaviors were tested, which storage and network assumptions were priced, which failure paths were exercised, and which ownership boundaries were approved before a purchase order is signed.
1What a procurement checklist means in a diskless Kafka design
Procurement becomes an architecture review when durable storage is separated from broker compute. In a traditional Shared Nothing architecture, each broker owns local log data and replication keeps copies on other brokers. A buyer can compare instance and disk prices, but the real operating cost also includes partition movement, recovery bandwidth, local-disk headroom, and the people required to keep those pieces healthy.
An object-storage-backed design changes the questions. A team still needs to know the write acknowledgment path, cache policy, metadata store, object request pattern, and network route. It also needs to know which Kafka semantics remain intact for producers, consumers, transactions, Consumer groups, offsets, Kafka Connect, and operational tooling. “Uses S3” is a storage detail; it is not a production acceptance criterion.
Start the review with a written workload contract. Capture the topics and partitions that matter, the expected write and read shape, retention and replay behavior, latency objectives, availability boundaries, security requirements, and growth assumptions. The values can be ranges where the workload is uncertain. What matters is that a vendor must test against the same contract that the platform team will use for the decision.
A diskless Kafka purchase is ready for approval when every important promise has an owner, a measurement, and a failure response.
The first worksheet can stay vendor-neutral. It should force the buyer to ask for evidence before debating product features.
| Procurement question | Evidence to request | Decision gate |
|---|---|---|
| Where does a record become durable? | A write-path diagram and a test showing acknowledgment behavior during storage latency or interruption | The acknowledgment boundary is explicit, and the recovery source is identified |
| What remains on the broker? | Cache, WAL (Write-Ahead Log), metadata, and local-retention settings for the proposed deployment | Local media has a defined role, capacity limit, and failure response |
| Who owns durable data? | Object-storage account, bucket, encryption key, lifecycle, and access-control ownership | Security and platform teams agree on the data boundary |
| How are reads served? | Tailing Read and Catch-up Read measurements under the target consumer pattern | Cache misses and historical reads have acceptable behavior |
| How does a broker leave? | Broker replacement, scale-in, partition reassignment, and rollback runbooks exercised on representative data | Recovery does not depend on an undocumented manual copy |
| What does the bill include? | Compute, object capacity, object requests, network transfer, private connectivity, observability, and support assumptions | The model can be reconciled with provider invoices and workload growth |
| Which Kafka contracts are preserved? | Client, security, transaction, offset, Connect, and monitoring test results | Application owners sign off on compatibility and migration scope |
The checklist is more useful when a failed gate has a named next action. “Catch-up reads are slow” should lead to a cache or workload test, not a vague request for better performance. “The bucket policy is unclear” should lead to an owner and an access review, not a promise that the provider handles it.
2The mechanism: brokers, cache, metadata, and object storage
The storage path explains why the checklist needs more than a price quote. In a Shared Nothing Kafka cluster, a partition leader appends to broker-local storage, followers replicate the data, and consumers fetch from brokers that own the relevant replicas. Scaling or replacing a broker therefore involves both compute and data placement.
Diskless Kafka moves the durable user-data boundary toward shared object storage, but the hot path still needs fast handling. A typical design uses a local cache for data that is being written or read frequently, a WAL for durable acknowledgment and recovery, object storage for the long-lived stream, and a metadata layer for offsets, ownership, and coordination. The exact implementation varies, so the buyer should ask the candidate system to map each responsibility to a component and a test.
The distinction between cache and source of truth is critical. Cache contents can be evicted and rebuilt. Durable stream data cannot be treated as disposable only because it is remote. During an outage, the team must know whether a replacement broker rehydrates from object storage, replays a WAL, follows metadata ownership, or performs some combination of those actions.
Reads also split into different operating paths. Tailing reads follow the head of a stream and usually depend on memory, local working space, and the write path. Catch-up reads reach older data and may require object retrieval and prefetch. A procurement test that measures only steady-state producer latency can miss the behavior that matters during replay, incident recovery, and backfill.
Network placement belongs in the same diagram. Producer and consumer traffic, broker-to-object-storage requests, metadata traffic, and observability exports can follow different routes. The cost and failure behavior depend on region, Availability Zone (AZ), private endpoints, routing policy, and the storage service selected. Ask for the route of each path and price the path that actually crosses a boundary; do not hide all traffic in a single “network” line.
Apache Kafka's Tiered Storage documentation is a useful comparison point. Tiered Storage moves eligible log segments to a remote tier while retaining a local tier on brokers. A diskless Kafka design can go further by making shared storage part of the primary durable model. The procurement question is not which label sounds better; it is whether the proposed system changes the broker, storage, and recovery responsibilities your team is paying to operate.
3Failure, cost, and compatibility checks
Failure testing should begin before a purchase is treated as a capacity exercise. The table below turns common claims into observable checks.
| Failure or operating event | What to measure | What a passing result should explain |
|---|---|---|
| Broker process or node loss | Producer acknowledgments, consumer progress, ownership handoff, and recovery traffic | Which data is already durable, which state is rebuilt, and when clients resume normal behavior |
| Object-storage latency or temporary errors | Write backpressure, retry behavior, cache use, and alerting | How the system protects ordering and durability without silently dropping records |
| Network isolation from the storage endpoint | Client-visible errors, fencing, retries, and operator actions | Whether the cluster fails closed, queues work, or serves a defined degraded mode |
| Scale-out or scale-in | Time to add or remove compute, partition movement, cache warmup, and network change | Whether elasticity requires copying retained data or changing ownership and metadata |
| Historical replay or catch-up | Fetch latency, object requests, cache eviction, and impact on tailing consumers | How cold reads share resources with real-time traffic |
| Configuration or rollout rollback | Data visibility, offsets, ACLs, and client reconnection | How the team returns to the previous state without rebuilding the log |
Cost review needs the same discipline. Object capacity is only one meter. Include request volume, retrieval patterns, cross-AZ or inter-region transfer, private networking, local cache and WAL media, broker compute, observability, backup, support, and migration overlap. Use the provider's current pricing pages for the region under review, write down whether each figure is list or negotiated pricing, and keep workload assumptions beside the formula. If the design uses a different WAL type for lower latency or a different durability boundary, model that choice separately.
Compatibility is a procurement gate because Kafka applications rarely stop at produce and consume. Test the client versions and producer settings used in production, including idempotent and transactional producers where applicable. Exercise Consumer groups, offset commits, retention changes, log compaction if used, security protocols, ACLs, Kafka Connect jobs, Schema Registry integration, and the monitoring signals that on-call teams rely on. The test should include a rollback path so that a green migration result does not hide an unsafe exit.
The security review follows the bytes. Confirm which identity can write to object storage, which identity can read it, how encryption keys are owned and rotated, how retention and deletion policies are enforced, and how audit records are exported. A customer-controlled deployment may require the data plane and storage bucket to remain inside a specific Virtual Private Cloud (VPC) or account. Those are purchase requirements, not implementation details to settle after the contract is signed.
4How AutoMQ changes the operating model
Once the checklist points to durable data that should not be tied to broker lifecycle, a Shared Storage architecture becomes a concrete category to evaluate. AutoMQ is a Kafka-compatible, cloud-native streaming platform in that category. It keeps the Kafka protocol and ecosystem contract while using S3Stream and S3-compatible object storage as the durable storage foundation.
That architecture changes what the buyer measures. AutoMQ Brokers handle Kafka requests, partition leadership, coordination, and cache behavior, while persistent stream data is organized through S3Stream. WAL storage sits in the write path to protect low-latency durability and recovery before data is uploaded and organized in S3 storage. Data caching serves hot and prefetched data, but the cache is not the long-term source of truth.
The WAL choice belongs in the acceptance record. AutoMQ Open Source uses S3 WAL, while AutoMQ BYOC (Bring Your Own Cloud) and AutoMQ Software can expose additional WAL storage choices according to the deployment and product configuration. The buyer should tie each choice to a latency target, durability boundary, failure domain, and cost model rather than treating “WAL” as one interchangeable performance number. The AutoMQ WAL storage documentation describes the role of the layer and the available options.
Stateless brokers are useful only when the storage and metadata contracts behind them are clear. If a broker can be replaced without copying its retained log from an attached volume, scaling and recovery can focus more on ownership, cache warmup, and traffic placement. That does not remove testing: the team still needs to verify fencing, replay, object-storage backpressure, metadata recovery, and client behavior under failure.
Deployment boundaries can be part of the buying decision as well. AutoMQ BYOC places the control plane and data plane within the customer's cloud environment, while AutoMQ Software targets customer-owned private data centers. Teams with account, VPC, encryption, or audit requirements should validate those boundaries with the same care as protocol compatibility. The relevant question is where the records, control actions, and operational telemetry travel, and who can authorize each path.
AutoMQ is not a shortcut around workload validation. A small workload with short retention may favor a different service model. A mature team with automation built around broker-local Kafka may decide that Tiered Storage addresses its immediate pressure. AutoMQ is a candidate when Kafka compatibility, independent storage growth, elastic broker operations, and customer-controlled deployment need to be evaluated together.
5Decision checklist and FAQ
Before approving a diskless Kafka purchase, ask the review group to sign the following gates:
- Workload fit: The write, read, retention, replay, and growth contract is written down and used in tests.
- Durability path: The team can point to the acknowledgment boundary, WAL behavior, object-storage source of truth, and metadata owner.
- Read behavior: Tailing and catch-up reads are measured separately, including cache misses and replay contention.
- Failure recovery: Broker, storage, network, and rollback drills have owners, evidence, and a documented client-visible outcome.
- Cost model: Compute, local media, object storage, requests, network, observability, support, and migration overlap are priced for the selected region.
- Compatibility: Kafka clients, Consumer groups, offsets, transactions, Connect, security, and monitoring pass workload-specific tests.
- Governance: IAM, encryption, bucket lifecycle, audit, region placement, and deployment boundaries have explicit owners.
- Rollout plan: A pilot, cutover, rollback, and post-cutover observation window are budgeted and scheduled.
5.1Is diskless Kafka the same as Apache Kafka Tiered Storage?
No. Tiered Storage adds a remote tier for eligible log segments while keeping an active local tier on brokers. Diskless Kafka usually describes a deeper change in which object storage participates in the primary durable model and local disks serve cache or working-storage roles. Verify the write path, local-retention role, and recovery unit for the system being evaluated.
5.2Does “diskless” mean that a cluster has no disks?
No. A diskless design may still use local media for cache, WAL, metadata, or short-lived buffering. The procurement question is whether broker disks remain the primary durable owner of user data and whether a failed broker must be recovered from its attached volume.
5.3Which numbers should a vendor provide?
Ask for measurements under your workload contract rather than generic headline figures. Useful evidence includes write acknowledgment behavior, tailing and catch-up reads, object request patterns, recovery traffic, scale-out and scale-in behavior, and the cost assumptions used to produce the quote. Every number should include its workload, region, storage choice, and test date.
5.4Where does AutoMQ fit?
AutoMQ fits the Shared Storage architecture branch of the checklist. It preserves Kafka compatibility while using S3Stream, WAL storage, data caching, and S3-compatible object storage to separate broker compute from durable stream storage. The decision still depends on the compatibility, failure, governance, and cost gates above.
The checklist starts with a simple question: what exactly are you buying when you buy a diskless Kafka platform? By the end of the review, the answer should be more precise than “Kafka on S3.” It should identify the durability boundary, the cache and metadata responsibilities, the failure behavior, the billable paths, and the tests that let application owners sign off.
If those gates point toward a Kafka-compatible Shared Storage architecture, use the same workload contract for a hands-on evaluation. Start with AutoMQ Cloud, then carry the compatibility, recovery, and cost checklist into the pilot.
