Blog

Reading Kafka Vendor Benchmarks with Skepticism: The Controlled Comparisons That Matter

Table of Contents

Table of Contents

A vendor benchmark lands in a planning meeting with a compound claim. The summary slide shows the platform under test sustaining roughly three times the throughput of the current cluster, at a tighter latency percentile, for less money. The bars are clean and the axes look honest. A preference starts to form before anyone has asked what the bars were fed.

Then a platform engineer tries to reproduce it. Same cloud, same client count, roughly the same instance family, and the workload ported as faithfully as the slide allows. The triple does not appear. Nothing fails; the number is not there. The team is left with three explanations: the vendor falsified the chart, the reproduction was badly run, or the claim was true only inside a laboratory-shaped universe. In practice the third explanation covers most cases.

That gap between a laboratory and a production environment is what Kafka benchmark skepticism exists to close. The chart was usually not faked. It was the honest output of a test whose conditions were never your conditions. Vendor benchmarks are marketing until the methodology holds up, and the useful skill is deciding which parts of the chart survive contact with your own workload.

1Numbers that vanish when you try to repeat them

The first instinct after a failed reproduction is to blame the harness. The second is to blame honesty. Both miss the more common mechanism. A Kafka benchmark is not a property of a platform; it is the product of a platform and a set of held variables. Change the variables and the result moves, sometimes enough to reverse the conclusion.

That is why the same chart can be true in a laboratory and invisible in production. The laboratory fixed the message size, the batch behavior, the acknowledgement mode, the partition count, the retention window, the read position, and the hardware. Production lets every one of those move. The vendor did not lie about the software; it reported a point in a design space that you do not actually occupy.

Reproduction usually fails quietly, which makes it worse. There is no error in the logs and no bug ticket to file. There is only a number that should be there and is not, while a procurement deadline asks whether the platform is slower than promised or the test was wrong. The honest answer is usually neither. The test measured something else.

So before doubting the platform, check the conditions. Most unrepeatable Kafka benchmarks fail on one of three axes, and the next two sections name them. If the conditions are not written down anywhere, that absence is itself the finding.

2The flaw list: configs, shapes, baselines

Three flaws show up again and again in vendor benchmark decks, and each one is easy to miss because it hides inside a chart that looks complete.

  • Configuration drift. Distributions ship different defaults for producer acknowledgements, replication, log pre-allocation, compaction, and throttling. One system measured with acks=1 next to another measured with acks=all is not an engine comparison; it is a durability comparison wearing a throughput costume. Ask for the configuration file, not the chart.
  • Load shapes that flatter one side. Steady-state writes of large compressed batches stress a different bottleneck than small messages with strict tail latency, and neither resembles a bursty profile with lagging consumers. A benchmark that never names its load shape is measuring an unknown quantity and drawing a line through it.
  • The missing cost baseline. Throughput per broker means little until the test pins what each result costs: storage per GiB, replication traffic, cross-zone transfer, and deliberately idle capacity. A chart that omits cost lets an expensive setup hide behind a quick bar.

Map of three benchmark flaws: configuration drift, mismatched load shape, and missing cost baseline

None of these flaws requires dishonesty. A vendor can publish a chart that is arithmetically correct and still answer a question you did not ask, because the configuration was favorable, the load shape flattered the architecture, and the cost baseline was left off the page. The flaw splits in two: the benchmark itself, and the reader who accepts it.

The cure is not to distrust all benchmarks. It is to demand the conditions, and then to compare systems as if conditions were part of the result rather than footnotes. That demand is what the four principles put into practice.

3Four principles of a fair comparison

A controlled comparison pins the variables the flaws let drift. The four that matter most are load, cost, data layout, and latency target. If two systems are not equal on all four, the comparison is measuring the variable it forgot, not the platform it thinks it is testing.

  • Same load. Replicate the message size distribution, key distribution, producer and consumer rates, partition count, and retention before any numbers are recorded. A benchmark without a written load profile is a screenshot of a moment, not a result.
  • Same cost. Fix the spend per unit of work on both sides: instance types, storage class, network placement, and reserved or idle capacity. Comparing a provisioned cluster against a minimally sized one proves that bigger budgets buy more throughput, which everyone already knew.
  • Same data layout. Partitions, replica placement, and the split between hot and cold reads change where the work goes. Apache Kafka's replication documentation describes the movement that a layout change can introduce; a comparison that ignores layout is comparing two different clusters with the same logo.
  • Same latency target. Pick the service level that matters, such as P99 produce latency under acks=all, and test both systems up to that line instead of past it. Capacity beyond the target is interesting only if you can buy it at a price you already know. The min.insync.replicas configuration is the place to confirm what the broker-side durability contract means before you encode it in a pass line.

Four principles of a fair Kafka comparison: same load, same cost, same data layout, same latency target

A comparison that holds these four is rare, and rarity is the point. When you find a vendor chart that holds all four, treat it as credible evidence. When you do not, the missing principle tells you exactly which question the chart cannot answer, and you can stop before the slide does its work.

4Salvaging signal from vendor benchmarks

A flawed benchmark is not worthless. It is signal wrapped in noise, and the reading skill is to unwrap the part that survives. Every chart, even a flattering one, reveals something about the architecture it favors if you know which question to ask it.

If the benchmark showsExtract this signalCheck this first
Peak producer throughputApproximate ingest ceilingAcknowledgements, batching, compression, and test duration
A latency percentile curveTail behavior under that loadDurability settings, warm-up, and read position
Scaling or recovery timingThe operating model behind capacity changeWhether data movement is measured or discounted
Cost per workloadThe vendor's cost hypothesisWhether network, storage, and idle capacity are inside the number

The table turns a marketing slide back into an engineering artifact. Each row asks for the variable the chart is least likely to have controlled, and the answer is usually written in the methodology section or not written at all.

For a benchmark you plan to reproduce, write the conditions down the way a lab would. A minimal spec fits in a single YAML block and removes most of the ambiguity before the first message is produced:

yaml
benchmark: purpose: "decide whether platform B matches platform A under our write path" workload: record_size_bytes: 1024 key_distribution: "uniform" partitions: 100 producers: 20 consumers: 8 read_position: "tail" durability: acks: "all" min_insync_replicas: 2 replication_factor: 3 cost_envelope: instances: 6 storage: "same class both sides" network: "cross-zone transfer counted" success_line: metric: "P99 produce latency" target_ms: 30 while_throughput: "sustained, named"

A spec this short does not make the benchmark perfect. It makes the benchmark auditable, which is the property that most vendor charts deliberately or accidentally leave out.

5Designing the benchmark that answers your question

The last step is to stop benchmarking Kafka in general and start benchmarking the decision in front of you. Skepticism is not a reason to avoid vendor numbers; it is the reason to replace them with your own, built around the variable you actually care about.

A decision-shaped benchmark has a different shape from a marketing chart.

  • Start from a decision, not a platform: "can this system hold our write path under our cost model" is a question; "which system is faster" is not.
  • Pin the four principles before the first run, not after the results look wrong.
  • Measure the cost model as a first-class result, with network and storage as explicit line items rather than hidden defaults.
  • Write the pass line first. If the answer can only be read after the run, the test cannot fail honestly.
  • Record what was not measured, because the excluded variable is where the next surprise will come from.

Checklist for designing a Kafka benchmark that answers a specific decision

The same checklist applies when the vendor is AutoMQ. AutoMQ publishes its own benchmarks and architecture material, and they deserve the same treatment: hold cost and latency as constraints, not as conclusions, and check which conditions the published numbers fixed. What makes AutoMQ worth reading in this context is a different operating model rather than a different marketing posture. It is a Kafka-compatible cloud-native streaming platform built on a Shared Storage architecture, where durable stream data lives in object storage behind S3Stream, with a WAL (Write-Ahead Log) write path that keeps latency bounded on the ingest side.

That architecture moves the variables rather than making skepticism unnecessary. When durable data stops living on broker-local disks, the questions shift from replication bandwidth and partition movement to WAL behavior, object storage access patterns, read cache hit rates, and cross-zone data movement. A fair comparison still needs all four principles; the checklist points them at a different place.

Back to the planning meeting. The slide said three times, and the room heard a promise about production. What the room actually had was a hypothesis about one configuration, and the difference between the two is the entire job. Design the benchmark that answers your question, pin the conditions, and let the result have the last word. The AutoMQ open-source project is a practical place to get a reproducible baseline for the shared-storage rows of that test.

6References

7FAQ

7.1Is it dishonest when a vendor benchmark cannot be reproduced?

Not necessarily. Most unrepeatable results come from ordinary differences in configuration, load shape, and cost baseline rather than fabricated data. The chart describes a specific laboratory honestly; the problem is that the laboratory is not your production environment. Demand the conditions before you demand an explanation.

7.2Which flaw hides the most: configuration, load shape, or cost baseline?

Configuration drift is usually the first thing to check, because it is concrete and quickly verified from a configuration file. Cost baseline is the most consequential over time, since it decides whether a win is real once the bill arrives. Load shape sits in between and is the one most often left entirely unreported.

7.3Can I reuse a vendor benchmark as a baseline for my own proof of concept?

Only as a starting hypothesis, never as a target. Reuse the vendor's published conditions as input to your own workload profile, then change the variables that describe your traffic and your cost model. The value of a vendor benchmark is that it tells you what to reproduce, not what to expect.

7.4Where should I run a fair Kafka comparison?

The environment barely matters compared with the conditions. Stage, production-like, or a dedicated lab all work if the load shape, cost envelope, data layout, and latency target are pinned and the same on both sides. Favor an environment where you can control the clients and read the configuration of both systems.

7.5Does a Shared Storage architecture make benchmark skepticism unnecessary?

No. It moves the variables. A Shared Storage architecture takes durable data off the broker, so the checklist shifts toward the WAL write path, object storage access, read cache behavior, and cross-zone data movement. The four principles still apply; only the inspection points change.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.