Blog

Redpanda Benchmark Claims: How to Reproduce Kafka Performance Comparisons

Table of Contents

Table of Contents

A Redpanda benchmark headline can look decisive while leaving out the variables that decide whether the result applies to your Apache Kafka® workload. The same is true of a Kafka benchmark published by any vendor. Change the record shape, producer batching, durability setting, partition count, storage path, or consumer position, and the chart may answer a different question without changing a single label on its axis.

Vendor benchmarks can still be useful when the experiment is visible. Treat the experiment as the unit of evidence, then ask whether another team can rebuild the topology, run the same workload, collect the same metrics, and explain every difference that remains. The protocol below treats Redpanda as the Kafka-compatible streaming vendor named in the search query and keeps the comparison focused on Kafka clients, Kafka semantics, and Kafka operating conditions.

Use the protocol to find out which platform meets your workload contract under conditions you can reproduce. A universal winner is outside the scope of the experiment.

1A headline number is the end of a chain

Performance claims are often compressed into one sentence: a platform handled a given throughput, reached a particular latency, or used fewer resources. The sentence hides a chain of decisions. Someone selected a client and version, generated a record shape, chose a partition count, placed brokers and clients in a network, set acknowledgments and replication, warmed the cluster, and decided which samples to include. Every link changes what the final number means.

Start by asking what the claim actually measures. Producer throughput can be limited by client CPU, broker CPU, network bandwidth, storage acknowledgment, or the generator itself. Produce latency can refer to the time until a client receives an acknowledgment or to an end-to-end timestamp observed by a consumer. A consumer test may read tailing data, historical data, or a cache that was already primed. Those are different experiments even when they use the same topic.

The first reproducibility artifact should therefore be a one-page experiment contract. Write down the decision the benchmark must support, then list the variables that can change that decision:

  • Workload contract: record size, key distribution, compression, producer and consumer concurrency, traffic shape, retention, and expected read position.
  • Kafka contract: topic and partition layout, replication factor, acks, min.insync.replicas, idempotence, transactions, quotas, and security settings.
  • Resource contract: broker and client instance types, CPU architecture, memory, network limits, storage media, and placement across Availability Zones (AZs).
  • Measurement contract: warm-up and measurement boundaries, percentile definitions, error handling, metric sources, and the artifacts to publish.

This contract is more valuable than a large result table because it prevents a later comparison from quietly changing the question.

2Keep Redpanda and Kafka on the same experimental footing

A fair Redpanda Kafka benchmark starts with matched resources and matched Kafka behavior. Matching broker counts alone is not enough. If one side uses faster disks, more client machines, a different CPU architecture, or a topology that keeps clients close to brokers, the test is comparing resource envelopes rather than broker implementations.

Record the full topology before the first run. Include the number and type of broker and client nodes, storage class and capacity, network placement, security path, and any managed service limits. State whether a benchmark uses local SSD, block storage, network file storage, or object storage. If one platform has a vendor-specific storage mode, name it and keep that mode visible in the report.

The Kafka behavior must be matched as carefully as the hardware. Use the same client library and version where that is supported, or explain why the client differs. Apply the same record generator, partitioner, compression codec, acknowledgment level, and transaction mode. If a platform requires a setting to reach its intended durability, include that setting in both the experiment contract and the comparison narrative. A result obtained with acks=1 should never be presented as if it represented an acks=all workload.

Apache Kafka's producer configuration documentation is a useful source for the semantics of acks, batch.size, and linger.ms. Redpanda's Kafka client compatibility documentation provides the corresponding vendor context. Neither page supplies a workload-specific performance conclusion, which is why the conclusion has to come from the controlled run.

3Use a workload matrix instead of one synthetic stream

A single stream can reveal a bottleneck, but it cannot represent the operating profile of most Kafka platforms. Build a small matrix where each scenario answers a different production question. Keep the rows stable across Redpanda, Apache Kafka, and any other Kafka-compatible system under review.

ScenarioWhat it isolatesMinimum evidence to save
Sustained produceWrite path under the declared durability contractProducer throughput, acknowledgment latency percentiles, error rate, broker CPU, storage and network utilization
Tailing consumeFresh-data delivery and consumer schedulingConsumer throughput, end-to-end latency percentiles, consumer lag, fetch errors
Catch-up consumeHistorical reads and interference with active writesRead throughput, lag recovery, write-path impact, storage reads, network utilization
Burst or fan-inHeadroom when many producers arrive togetherQueue time, request latency, throttling, rejected requests, and recovery to steady state
Failure and recoveryBehavior during broker or storage disruptionFailure event, unavailable partitions, lag growth, recovery time, manual actions, and final state

The matrix keeps each test focused enough to answer a production question. That prevents a peak producer run from being mistaken for a complete Kafka performance comparison. A platform can lead a write-only run and still miss the requirement that matters most to a team: predictable tail latency while consumers replay retained data, or a bounded recovery path after a broker is removed.

Use the same topic shape and the same data generator in each scenario. If a scenario intentionally changes partition count, record that change as part of the question rather than presenting the result as a direct platform comparison. The test harness should emit a manifest with the configuration, environment, source revision, and run identifier so that a result can be traced back to the exact workload.

4Report tail latency as a distribution

Average latency is convenient to read and prone to misuse. It can remain stable while a small portion of requests waits behind storage, replication, garbage collection, or a client retry. Those requests are the ones that often violate an application SLO. A Kafka latency benchmark should publish P50, P95, P99, and, where the workload requires it, a higher tail percentile, together with the measurement window and sample policy.

Separate the latency layers instead of collapsing them into one line:

  • Producer acknowledgment latency measures the client request path under the declared acks and batching settings.
  • End-to-end latency measures the event timestamp to consumer observation, including broker, storage, replication, fetch, and client scheduling.
  • Recovery latency measures how the distribution changes during a declared failure or reassignment event.
  • Queue and resource signals explain whether the tail is caused by client backpressure, broker queues, storage wait, network saturation, or consumer scheduling.

The distribution is the result. The resource signals are the explanation. Save both. Without them, a reproduction can show that a number changed but not why it changed, which makes tuning and review guesswork.

A benchmark report should also state how outliers are treated. Do not silently discard timeout samples, failed requests, or warm-up records. If an experiment excludes startup or ramp traffic, mark the boundary in the result file and retain the excluded samples. The protocol should make it possible to rerun the same decision with a different outlier policy without losing the original evidence.

5Make durability and replication visible

Durability is part of performance. acks, replication factor, min.insync.replicas, idempotence, and transactions change the work required before a producer receives an acknowledgment. A benchmark that changes these values to improve one platform's chart has changed the durability contract.

Declare the durability profile beside every result. The profile should answer what a successful write means, how many replicas or storage acknowledgments are required, and what happens when a replica falls behind. If the platform uses a different persistence architecture, describe the equivalent contract in operational terms. The comparison remains fair when the reader can see what each system promises before the acknowledgment is returned.

This is also where failure tests become useful. Stop or isolate the declared failure target, observe the producer and consumer behavior, and record how the platform returns to the healthy state. For broker-local Kafka, replica catch-up and partition movement may dominate the recovery path. For a shared-storage Kafka-compatible system, metadata ownership, WAL replay, cache refill, and storage dependency may dominate. The test should measure the architecture's real work rather than assuming that every system fails in the same way.

Keep the fault injection deterministic. Use a scripted action, record its start and end timestamps, and capture the cluster state before and after. If the operator had to intervene, include the exact command or runbook step in the artifact. “Recovered” is incomplete unless the report says whether producers stayed available, how consumer lag behaved, and what manual action made recovery possible.

6Add cost only after the performance contract is fixed

Cost comparisons are useful when they describe the resources consumed by the same workload. They are misleading when they turn one benchmark number into a generalized price claim. Keep performance and cost in separate result sections, then connect them through a declared resource and pricing model.

A reproducible cost model should identify at least:

  • broker and controller compute;
  • storage capacity and storage operations;
  • inter-AZ and cross-region network paths;
  • client, connector, and observability resources;
  • license or service charges; and
  • the measurement period and cloud region used for the price inputs.

Do not publish a dollar figure unless the inputs are current and cited in the review record. Prices vary by provider, region, contract, and date. A benchmark can still report resource consumption and leave the monetary conversion to the reader's current price sheet. That is more useful than a precise cost claim whose assumptions have already expired.

Cost per unit of useful work is a better comparison than cost per broker. Define the unit first, such as retained GiB, successfully processed records, or an end-to-end event volume that includes the required durability and replay scenarios. Then show the formula and each input. If a platform changes the storage or replication model, explain that architectural difference next to the formula rather than hiding it inside a blended number.

7Include AutoMQ as a controlled test subject

Once the protocol is fixed, the experiment can include AutoMQ under the same workload contract. AutoMQ is a Kafka-compatible cloud-native streaming platform with a Shared Storage architecture, so its storage path and broker behavior need to be described rather than treated as an implementation detail. The fair comparison is to keep the client, workload, durability question, and reporting format constant, then let each architecture show where it spends work.

For AutoMQ, record the storage mode and deployment boundary as part of the topology. AutoMQ Open Source uses S3 WAL; AutoMQ BYOC and AutoMQ Software can use the WAL type supported by the chosen deployment. Do not mix those modes in one result row. The AutoMQ architecture overview explains the distinction between Kafka compute, WAL storage, data caching, and S3 storage.

The same discipline applies to the interpretation. A shared-storage design changes the cost and recovery questions because durable data is not tied to one broker's local disk. The report should therefore include storage dependency, cache behavior, scaling events, and recovery evidence so a buyer can judge whether the architecture fits the workload.

Readers who need a generic Kafka test plan can use the existing fair comparison methodology as a companion. The benchmark interpretation guide covers the buyer questions that remain after the raw results are collected. These links support the process; they do not replace a run against the versions and topology under review.

8A runbook that another team can execute

The benchmark should be stored as code and configuration, not as a slide deck. Keep the harness, environment manifest, platform configuration, topic setup, client properties, fault scripts, metric queries, and result parser in one versioned bundle. Pin the source revision for every component. A later rerun should be able to identify whether a result changed because of code, configuration, infrastructure, or a platform release.

A minimal producer run can use the Apache Kafka performance tool with variables supplied by the experiment manifest:

bash
kafka-producer-perf-test.sh \ --topic "$TOPIC" \ --num-records "$NUM_RECORDS" \ --record-size "$RECORD_SIZE" \ --throughput "$THROUGHPUT" \ --producer.config producer.properties

The command is a harness, not a benchmark conclusion. Keep producer.properties, the topic configuration, and the platform manifest beside its output. For consumer tests, use the corresponding consumer tool or the production client with the same deserializer, isolation level, and group behavior that the application uses. A production client test is often the right choice when the workload depends on transactions, custom serializers, or application-level timing.

Before the measurement window, run a preflight that proves the client generators and network can drive the declared load. During the run, collect client and platform metrics on the same clock. After the run, archive raw samples, logs, configuration, and a machine-readable summary. Have a second engineer recreate the environment from the bundle before treating a result as a procurement input.

9What a trustworthy result looks like

A credible Redpanda performance comparison does not hide behind a single “faster” label. It shows the workload contract, matched resources, Kafka durability semantics, latency distribution, failure behavior, and cost boundary. It states what the test did not cover. It keeps vendor claims separate from measured results and marks any unverified claim as a question for a fresh run.

When a Redpanda benchmark claim cannot be reproduced, the result is still useful if the missing variable is documented. Perhaps the vendor used a storage mode that is not available in your region. Perhaps the client count, compression ratio, or partition layout was omitted. Those are findings about decision risk, not invitations to fill gaps with assumptions.

Return to the opening chart with the experiment contract beside it. The chart may still be impressive, but now you can tell which workload it represents, which Kafka guarantees it includes, and which production questions remain unanswered. If AutoMQ is on the shortlist, start with the AutoMQ GitHub repository, reproduce the same contract, and bring the measured evidence into the architecture review.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.