Table of Contents
Table of Contents
Redpanda pricing is easy to search for and surprisingly difficult to model. A broker or cluster price does not tell you what an Apache Kafka workload will cost after retention, replication, consumer fan-out, failure headroom, and recovery traffic are included. Those multipliers are created by the workload and the storage architecture, so a price-page comparison can be precise and still answer the wrong question.
The useful question is: how many bytes does this Kafka workload make the platform store, copy, move, and keep available? This article builds a reusable model for Redpanda and other Kafka-compatible platforms. It uses explicit assumptions instead of Redpanda price claims. Replace the assumptions with a dated region, service tier, and cloud price sheet before making a purchasing decision.
1Start with logical data, not a broker count
The first input is the logical stream entering Apache Kafka. Let I be average ingress in GiB per day and H be the retention window in days. If compaction is not reducing the retained set, the logical retained data is:
L = I × H
L is the data that producers and consumers think they are retaining. It is not yet the physical storage requirement. Kafka replication stores multiple copies, and the broker also needs room for segment indexes, metadata, compaction overhead where applicable, rolling operations, and bursts above the average rate.
That distinction matters for Redpanda because local-storage capacity, instance selection, and Tiered Storage policy can be connected. Redpanda's Kafka client guidance and topic configuration reference are useful starting points for inventorying the workload, but they do not turn an average ingress number into a capacity plan. The topics, partitions, retention overrides, and consumer groups still have to be measured in your environment.
Consider a deliberately small illustrative workload:
| Input | Assumed value | Why it is an assumption |
|---|---|---|
Average logical ingress (I) | 500 GiB/day | A planning value, not a Redpanda measurement |
Retention (H) | 7 days | A policy choice that must be confirmed by application owners |
Replication factor (K) | 3 | A durability choice; verify per topic and environment |
Storage overhead (O) | 15% | A planning allowance for indexes, metadata, and segment behavior |
Target physical utilization (U) | 70% | A headroom policy for bursts and broker operations |
With these assumptions, L = 500 × 7 = 3,500 GiB. The calculation is intentionally transparent. Change the retention window or ingress rate and the result changes without pretending that the vendor has published a universal number.
2Replication turns retained bytes into physical capacity
For a topic with replication factor K, a first-order physical storage estimate is:
S_physical = L × K × (1 + O)
S_reserved = S_physical ÷ U
Using the illustrative inputs, S_physical = 3,500 × 3 × 1.15 = 12,075 GiB, and S_reserved = 12,075 ÷ 0.70 ≈ 17,250 GiB. These are planning outputs. They do not describe a Redpanda cluster size, a Redpanda Cloud entitlement, or a cloud invoice.
The replication factor is not a decorative setting in a cost model. It affects storage, recovery work, and often network traffic. For new records, the leader sends copies to followers in the in-sync replica set. A simplified logical replication volume is:
R_replication ≈ I × (K - 1)
The approximation assumes that every record is replicated to K - 1 followers and that protocol overhead is handled in the storage overhead allowance. Real traffic depends on batching, compression, acknowledgments, placement, recovery, and whether replicas cross availability zones. Use measurements from broker metrics when replacing this approximation with an operating budget.
Replication can also create temporary cost spikes. A broker replacement, partition movement, or failed replica may copy a large fraction of the retained set while the workload continues to ingest and serve consumers. Budgeting only the steady-state copy rate hides the capacity and traffic needed to recover from an ordinary failure.
3Retention is a multiplier on every replay and recovery path
Retention is often treated as a storage setting. In a Kafka system it is also a recovery and network setting. A retained segment may be read by a replay consumer, fetched by a new replica, copied into a backup, or served to a second region. The longer the retention window, the larger the set of bytes that can participate in one of those paths.
Separate at least four byte flows in the worksheet:
- Ingress bytes are records accepted from producers before replication.
- Replica bytes are copies sent between brokers to satisfy the replication factor.
- Consumer bytes are records fetched by each consumer group, multiplied by the groups that read the topic.
- Recovery and backup bytes are moved during replica rebuilds, restore tests, snapshots, or planned migrations.
The flows are related, but they are not interchangeable. Counting consumer fan-out as retained storage overstates disk requirements. Omitting fan-out from network planning understates traffic. Combining both into a single “throughput” line makes it impossible to explain why two clusters with the same producer rate receive different bills.
For each flow, add a placement factor X for the fraction that crosses a chargeable boundary, such as an Availability Zone (AZ) or region:
N_chargeable = N_total × X
C_network = N_chargeable × P_network
P_network is a price input to verify from the selected cloud and region. Leave it blank until the pricing date and direction are known. Cross-zone transfer, cross-region transfer, service endpoints, and managed-service network policies can all use different rules. Do not copy a rate from an older calculator or from a different region into the Redpanda model.
The existing Kafka cross-AZ cost analysis is a useful companion when the placement factor is hard to estimate. Expose the cross-zone fraction instead of assuming that every byte crosses a zone, then use the same worksheet to evaluate a topology change.
4Compute and idle capacity belong in the same model
Storage and traffic explain only part of the spend. Kafka-compatible brokers are provisioned for CPU, memory, connections, partition leadership, and peak request rates. A retention-heavy workload can therefore pay for compute capacity selected to obtain local disk or network throughput. A bursty workload can pay for peak capacity that is idle during most of the observation window.
Use a separate compute term rather than folding it into storage:
C_compute = Σ (broker_hours × verified_instance_rate)
C_service = verified_vendor_or_managed_service_charge
C_storage = Σ (physical_storage_units × verified_storage_rate)
The rate inputs must be tied to a cloud, region, instance family, service tier, and date. The Redpanda Cloud billing documentation describes billing dimensions that should be checked against the selected service mode. It does not justify inserting a remembered dollar value into a new estimate. For a self-managed or BYOC design, keep the vendor service charge and customer cloud bill as separate rows.
Headroom deserves its own sensitivity check. If the cluster is sized at an average rate and runs close to the utilization target, one traffic burst or one broker drain can trigger a resize or a rebalance. If it is sized for the peak, the unused capacity becomes part of the monthly cost. Model at least three workload bands:
| Band | Inputs to vary | What the result tells you |
|---|---|---|
| Baseline | Average ingress, normal consumer groups, policy retention | The recurring cost floor under normal operation |
| Peak | Peak ingress, burst duration, consumer catch-up rate | Whether peak capacity or autoscaling drives the bill |
| Recovery | Replica rebuild volume, broker drain, replay traffic | The cost and capacity needed to return to a healthy state |
The same broker count can look efficient in the baseline and expensive in recovery. A model that reports only the average hides that difference.
5A scenario worksheet that can be audited
Create one row per topic class or workload rather than one row for the whole cluster. A compact worksheet needs the following columns:
| Column | Formula or input | Verification question |
|---|---|---|
| Logical ingress | GiB/day | Is this producer payload before or after compression? |
| Retention | days or GiB | Is the policy topic-level, and does compaction change the retained set? |
| Replication factor | K | Is the value the same for all topics and environments? |
| Physical overhead | O | Is the allowance based on measurements or explicitly conservative? |
| Utilization target | U | Does it leave room for bursts and broker replacement? |
| Consumer fan-out | groups and GiB/day | Which consumers read across a zone or region boundary? |
| Recovery volume | GiB/event and events/month | How often are rebuilds, restores, or migrations rehearsed? |
| Price inputs | provider and vendor rates | Are all values dated, regional, and tied to a service tier? |
Then calculate the billing-period view using a stated conversion. If the worksheet uses D days per billing period, do not silently assume a calendar month. Keep fixed, hourly, and daily rows separate so that a daily traffic estimate is not multiplied twice:
C_period = C_service_period + C_compute_period + C_storage_period
+ C_network_period + C_backup_recovery_period
+ C_operations_period
C_daily_row_for_period = C_daily_row × D
C_operations_period can be an internal labor estimate or a separately tracked run-rate. If it is excluded, label the result “infrastructure-only” instead of presenting it as TCO. Procurement teams can then compare vendor charges and cloud infrastructure without losing the operational boundary.
The worksheet should produce sensitivity views, not one magic number. Vary retention, ingress, replication factor, cross-zone fraction, and peak-to-average ratio. The most useful output is often a break-even boundary: for example, the retention window at which a storage architecture with more local capacity becomes less attractive than one with shared durable storage. The boundary depends on verified price inputs; the formula remains useful when those prices change.
6Where shared storage changes the cost equation
The workload model is useful even when the answer is to keep Redpanda. It also reveals when the central cost driver is the coupling between durable data and broker-local capacity. If retained bytes grow independently of compute, or if scale-in requires moving partition data between brokers, changing a broker size may not address the underlying cost.
That is the point at which a different architecture deserves evaluation. AutoMQ is a Kafka-compatible streaming platform with a Shared Storage architecture: durable data is stored in shared object storage, while Stateless Brokers provide the compute path. This does not remove cost; it moves the cost model to compute, WAL (Write-Ahead Log) storage, object storage, requests, and network paths that must be priced for the chosen deployment.
The comparison should preserve the same workload rows and change only the architecture-specific terms:
| Model term | Broker-local storage model | Shared Storage model |
|---|---|---|
| Retained data | Local or tiered storage tied to broker capacity | Shared object storage plus the selected WAL storage |
| Broker scaling | May require partition movement and data rebalance | Compute and durable data have a looser coupling; verify behavior for the selected deployment |
| Replication traffic | Broker-to-broker copies depend on replica placement | Durability path uses the shared-storage design; verify WAL and object-storage traffic |
| Price inputs | Broker instances, local or attached storage, service charge, network | Broker compute, WAL, object storage, requests, service charge, network |
| Main validation | Recovery and rebalance behavior under retained data | Read/write path, recovery, WAL choice, object-storage request volume, and compatibility |
The architecture does not make the decision on its own. Keep the Kafka contract, retention policy, consumer fan-out, and recovery objective constant, then compare the verified rows. AutoMQ's architecture overview and Stateless Brokers documentation describe the implementation boundary; the cost worksheet still has to be filled with your cloud and workload values.
7Decision gates for a Redpanda cost review
Before selecting a Redpanda service tier, a self-managed topology, or a Kafka-compatible alternative, require evidence for five gates:
- Retention gate: every topic class has an owner, a retention reason, and a measured retained-byte curve.
- Replication gate: the factor, placement, and rebuild behavior are documented for each production class.
- Traffic gate: producer, replica, consumer, backup, and recovery flows are separated, with chargeable boundaries marked.
- Capacity gate: baseline, peak, and recovery bands are modeled with utilization and headroom assumptions.
- Price gate: every dollar input has a provider, region, service tier, and verification date; unknown values remain
[VERIFY]until checked.
If a cluster fails the first four gates, a price comparison is premature. If it passes them and the price gate is still sensitive to retention, broker-local capacity, or cross-zone replication, run the same workload through a second architecture and compare the assumptions. That keeps the decision technical and auditable.
8FAQ
8.1Does Redpanda cost less than Apache Kafka?
There is no universal answer. Redpanda, self-managed Apache Kafka, and other Kafka-compatible platforms expose different service and infrastructure boundaries. Compare the same ingress, retention, replication, fan-out, recovery, and headroom assumptions with verified regional prices.
8.2How does retention change Redpanda cost?
Retention multiplies logical ingress into retained bytes. Replication and headroom then multiply that retained set into physical capacity. Longer retention can also enlarge replay, recovery, backup, and migration traffic, so storage-only estimates are incomplete.
8.3Is replication cost the same as storage cost?
No. Replication affects both stored copies and the bytes moved between brokers. Storage is a retained-capacity question; replication traffic is a placement and recovery question. Keep them as separate rows in the worksheet.
8.4Should I include Redpanda Cloud pricing in the model?
Yes, when the deployment uses Redpanda Cloud. Keep the service charge separate from the cloud infrastructure rows, and verify the billing dimensions and rates for the selected tier, region, and pricing date. For BYOC or self-managed deployments, include the customer-owned cloud bill explicitly.
8.5When should a team evaluate AutoMQ?
Evaluate it when Kafka compatibility remains important but the dominant cost driver is retained data tied to broker capacity, peak headroom, partition movement, or broker-to-broker replication. Use the same workload model, and verify AutoMQ compute, WAL, object storage, request, and network inputs for the target environment.
Retention is the first number in the worksheet, but it should not be the last number in the decision. If you want to test the model with a real workload, run an AutoMQ deployment assessment with your ingress, retention, replication, and recovery assumptions. The output is useful only when every price row is dated and every workload assumption has an owner.
