Blog

What 100 Percent Kafka Compatible Should Actually Require in a Proof of Concept

Table of Contents

Table of Contents

Sales decks like the phrase "100% Kafka compatible" because it compresses a hard technical question into one reassuring line. Engineering teams distrust the same phrase because it compresses a hard technical question into a slide nobody signed. A proof of concept is where the two meet, and most failed PoCs do not fail on the word Kafka. They fail on the word compatible.

Compatibility is not one property. It is a stack of at least five contracts that can break independently: the wire protocol, client runtime semantics, the ecosystem that attaches to a cluster, the schema and serialization layer, and the operational tooling a platform team already depends on. A hello-world producer and consumer exercise only the first two, and only their happy path.

Nothing in that smoke test tells you whether a Connect sink resumes from the right offset after a task restart, whether a transactional producer keeps its guarantees across a leadership change, or whether the existing Terraform modules, quotas, and dashboards survive the swap. The fix is not another vendor demo. It is a written plan that turns the claim into a finite set of executable checks, each with a pass line and a place on a timebox.

1Compatibility is five layers, not one slide

A "100% compatible" slide collapses distinct contracts into one phrase so that each gets credit for the others. Protocol compatibility can pass while transactional semantics fail. A connector can open a connection while its offset commits behave differently. A platform can be drop-in for producers and broken for the ops team that has to run it at 2 a.m. The proof of concept is the last cheap moment to separate those layers, before the decision gets expensive.

The five layers fail in different ways, which is why each one needs its own test and its own pass line.

  • Protocol. Produce, fetch, metadata, offset, and group coordination requests behave according to the Kafka wire format, so existing client libraries work without modification.
  • Client semantics. Idempotence, transactions, acknowledgement behavior, ordering, and consumer group rebalancing keep the same guarantees under retries and leadership changes, not only in the calm case.
  • Ecosystem. Connect source and sink connectors, Kafka Streams applications, Schema Registry clients, and third-party UIs exercise the API corners a hello-world program never touches.
  • Schema and serialization. The Schema Registry API and the wire-level SerDe round-trip work together, including evolution and compatibility checks.
  • Operations tooling. Admin operations, ACLs, metrics, quotas, and the runbooks built on them keep working so the team can operate the platform without building a parallel one.

Five-layer Kafka compatibility stack for a proof of concept, from protocol to client semantics, ecosystem, schema, and operations tooling

A layer diagram also fixes scope. If the vendor's compatibility matrix covers only the protocol layer, the PoC is not "80% compatible." It is one layer verified and four layers untested. Write that down before the meetings start.

2Protocol, clients, ecosystems, schemas, tooling

The Kafka protocol is the easiest layer to overestimate. A platform can serve produce and fetch requests correctly while diverging on less-traveled paths: version negotiation, error codes, metadata updates, or the wire behavior of group coordination. Protocol compatibility means the client library never needs to know which implementation is on the other end. The test is whether the version you actually deploy, not a demo build, handles your client matrix unchanged.

Client semantics are where "compatible" stops being a synonym for "connects." A producer with acks=all, idempotence, and transactions depends on broker behavior to guarantee that records are not duplicated or lost across retries and broker changes. A consumer depends on group coordination and offset commit behavior to resume cleanly after a rebalance. The delivery semantics and consumer configuration references describe the contract. The PoC has to exercise that contract under injected failure, not just measure that a record arrived once in a quiet environment.

The ecosystem layer exists because connectors and stream processors use and administer the cluster far more aggressively than a business application. They create topics, read group state, fetch offsets, and manage their own consumer groups. This is where subtle incompatibilities surface first. Kafka Streams state stores, Connect's offset and status topics, and third-party admin consoles all assume precise behavior. If a platform changes how offset commits or group states are exposed, the connector still runs but the operations team inherits the repair.

Schema and serialization is a separate gate because correct wire compatibility plus a broken Schema Registry still breaks every producer and consumer that shares typed records. Avro, Protobuf, and JSON Schemas wire compatibility tests verify payloads survive unchanged. The registry itself has a REST surface with subjects, versions, and compatibility modes.

Operations tooling is often the last layer considered and the first to fail the moment the PoC moves beyond a laptop. kafka-topics, kafka-consumer-groups, and ACL admin commands, plus JMX or Prometheus metrics, quotas, and the runbooks that depend on them, are de facto parts of the platform contract.

Keep that in view when a vendor leads with architecture instead of test results. The compute-storage-separation versus tiered-storage comparison is a useful companion, but only after the five compatibility layers earn their own pass lines.

3A minimal test case for each layer

Each layer gets one representative test that is small enough to run repeatedly and strict enough to surface a real difference. The pass line is behavioral, not architectural: a platform passes when the client, connector, schema, or operator cannot tell the difference, not when the vendor explains why the difference does not matter.

LayerMinimal test casePass line
ProtocolRun your production client matrix against produce, fetch, metadata, offsets, and group coordination with byte-identical payloads across many partitionsRequests succeed with expected offsets, ordering, error codes, and version negotiation; no client change required
Client semanticsEnable idempotence and transactions, then inject a leader change, broker restart, and duplicate sendIdempotent and transactional guarantees hold; retries produce no duplicates, gaps, or reordering beyond the delivery contract
EcosystemRun an existing Connect source and sink plus a stateful Kafka Streams job, then restart a task or nodeConnectors resume from the correct offset without duplication or loss; Streams state restores and reprocessing stays consistent
Schema and SerDeRegister a subject, produce and consume Avro records with a SerDe, then register a backward-compatible schema changePayloads round-trip intact; compatible changes are accepted and incompatible changes are rejected
Ops toolingExercise topic and group admin commands, ACLs, quotas, and metrics collection from the existing runbooksEvery documented admin action and metric works without a new tool or parallel workflow

Kafka compatibility PoC test matrix with one minimal test per layer and an explicit pass line

The table is the contract, not a checklist of features to admire. Run these five before any throughput or latency benchmark, because a platform that fails the semantic test should not spend another hour on a performance test.

4Timeboxing and pass lines for the PoC

A compatibility PoC does not need a quarter. It needs a sequenced plan that front-loads the cheapest decisive tests. Start with a two-day smoke gate, run the semantic and ecosystem work in the middle, and reserve the final days for failure injection and a written decision.

  1. Days 1 to 2: smoke gate. Protocol basics and one full produce-fetch round trip. Pass: your standard client connects, authenticates, and round-trips records. Fail and the PoC stops here.
  2. Days 3 to 7: client semantics. Run the production client matrix with idempotence, transactions, rebalances, and a live leader change. Pass: delivery guarantees hold across the injected events.
  3. Day 8 onward: ecosystem and schema. Run the real Connect pipelines, Streams jobs, and Schema Registry workflows. Pass: existing connectors and SerDes work without adaptation.
  4. Final three days: ops tooling and failure injection. Run the admin commands, quotas, metrics, broker replacement, and recovery drills. Pass: the on-call runbook completes with the same tools.
  5. Decision gate. A pass line is met only when every layer's test is green and any deviation is documented as accepted risk with an owner. A single unowned deviation is a fail.

Timebox plan for a Kafka compatibility proof of concept, with smoke gate, semantic tests, ecosystem checks, and a final decision gate

The timebox does more than schedule work. It forces the evaluation to call a pass or a kill by a fixed date instead of extending the PoC every time the vendor brings a new demo. Teams that extend the timebox without a new test plan are re-running the sales process, not the verification.

5Extra checks for shared-storage platforms

Some candidates claim Kafka compatibility while replacing the local broker storage model with a Shared Storage architecture. That changes what a durable write means, so the compatibility test needs to extend into the storage path rather than stop at the client boundary.

The added checks follow the data path, not the feature list:

  • Acknowledged write survives. Produce with acks=all, verify the record on the broker's durable layer, then remove the broker and return the record.
  • Upload completion is visible. Ensure data reaches object storage on a schedule that matches durability claims, and error or backpressure state is observable.
  • Older data reads correctly. Fetch records from hours and days back, not only the hot tail, and verify ordering, compaction, and retention behavior match the configured topic.
  • Recovery from storage, not replica. Replace a broker or force a partition movement and verify the leader resumes from the durable shared layer without guessing which local copy is current.

That last check exists because storage architecture changes the recovery path. Kafka's KIP-405 tiered storage design separates remote log data from local storage. A broader shared-storage approach also changes how leader and follower state behave, so recovery has to be measured against the new path.

That requirement is exactly where AutoMQ becomes a concrete candidate to test. AutoMQ is a Kafka-compatible platform that replaces broker-local replicated storage with a Shared Storage architecture built on S3Stream and S3-compatible object storage, with WAL storage serving the low-latency write path. Its compatibility with Apache Kafka documents the client surface, and the architecture overview and S3Stream overview describe the storage boundaries the extra checks target.

The point is not to test AutoMQ differently. It is to keep the extra storage checks from falling off the plan when the vendor page emphasizes protocol compatibility. The scaffold above works unchanged for AutoMQ, for Apache Kafka deployments, and for any Kafka-compatible platform that changes the storage layer.

6Move the claim off the slide

The next time a vendor deck says "100% Kafka compatible," ask for the test plan that backs the phrase. If the plan has one happy-path producer, one consumer, and a benchmark chart, the claim is still on the slide, not in the codebase.

Put the five-layer scaffold into the PoC instead. Test protocol, client semantics, ecosystems, schemas, and ops tooling with the real client matrix and real failure injection, add the storage checks when the architecture changes, and hold every deviation to a named owner and a decision date. That is what "compatible" should actually require.

If the storage layer is the open question in front of you, start from the AutoMQ repository on GitHub and apply the same five-layer test cases to your own workload.

7References

8FAQ

8.1What does "100% Kafka compatible" actually mean?

It means the protocol, client semantics, ecosystem connectors, schema and serialization, and operations tooling all behave like Apache Kafka from the consumer's point of view. Because each layer can pass or fail independently, the claim is only meaningful when a test plan exercises all five. A phrase in a deck is not a substitute for that plan.

8.2Which layers should a Kafka compatibility proof of concept cover?

Cover at least five: wire protocol, client runtime semantics, ecosystem plugins such as Kafka Connect and Kafka Streams, Schema Registry and SerDe behavior, and operations tooling including admin commands, metrics, and quotas. For platforms that change the storage architecture, add checks for durable writes, object storage upload, reading older data, and broker recovery.

8.3How long should a Kafka compatibility PoC take?

A focused plan can finish in two to four weeks. A two-day smoke gate decides protocol basics early, the middle weeks cover client semantics, ecosystems, schemas, and metrics, and the final days run failure injection and a written decision. The timebox matters less than a fixed pass line and a documented owner for every accepted deviation.

8.4Can a platform be protocol compatible but fail a PoC?

Yes. Protocol compatibility only covers the wire-level requests and responses. Transactional semantics, connector offset commits, Schema Registry behavior, and ops tooling can all diverge while a producer and consumer still exchange records. That is why the PoC needs a separate test and pass line per layer, with failures injected rather than observed only in a quiet run.

8.5What extra checks apply to shared-storage Kafka platforms?

A Shared Storage architecture changes where durable data lives, so the PoC should verify that an acknowledged write survives a broker failure, that upload to object storage is observable, that older data reads back correctly with compaction and retention intact, and that recovery starts from the durable shared layer rather than a local replica. These checks sit on top of, not instead of, the five standard layers.

8.6Do I need to benchmark performance before compatibility?

No. Run the compatibility tests first. A platform that fails a semantic or recovery test should not consume another hour of benchmark time. Once every layer passes, performance testing can run on the same client matrix with a clear workload, duration, and expected outcome.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.