Blog

Streaming data contracts on AutoMQ: compatibility has layers

Table of Contents

Table of Contents

“Kafka-compatible” answers an important question: can an existing Kafka client connect, produce, consume, and use the broker APIs it depends on? It does not answer whether a producer and a consumer agree on the meaning of every field they exchange. A platform can accept the same produce request while a schema change still breaks a downstream application.

That distinction matters when a team moves a workload to AutoMQ. The Kafka protocol remains the application boundary, while serialization, schema evolution, and business meaning remain application and platform concerns. The practical test is larger than a connection check: can the same registry, serializers, readers, tombstones, and rollback procedure work against the target data plane?

The useful design is layered. Let the Kafka-compatible platform move records, let a separate Schema Registry govern schema identity and evolution, and let producers and consumers enforce the parts of the contract that only they understand. Once those boundaries are explicit, an AutoMQ evaluation becomes a controlled compatibility exercise instead of a promise hidden inside a product label.

1Protocol compatibility is not contract compatibility

An Apache Kafka® Record has a transport shape: a key, a value, headers, a timestamp, and an Offset within a Topic Partition. The broker receives a request built by a client library and stores or serves the record according to Kafka semantics. The Apache Kafka protocol documentation describes that client-facing boundary. It does not define whether the value is Avro, Protobuf, JSON Schema, or an application-specific byte sequence.

The contract begins one layer above that boundary. A Producer serializes an application object, often after resolving a schema through a registry. A Consumer fetches the record, resolves the writer schema or wire format, and turns the bytes back into an application object. A registry can check whether a proposed schema is compatible with earlier versions, but it cannot determine whether status = "settled" still has the same business meaning after a product change.

That gives a streaming data contract at least four separate surfaces:

  • Transport: Topics, partitions, offsets, headers, acknowledgments, and Consumer group behavior.
  • Encoding: The serializer, deserializer, wire format, schema reference, and schema version carried or resolved with the record.
  • Evolution: The compatibility policy that decides whether a proposed writer and reader pair can coexist.
  • Meaning: Domain rules such as units, identity, allowed state transitions, and whether a missing value means unknown, empty, or deleted.

These surfaces interact, but they fail differently. A client can connect while a serializer cannot reach the registry. A schema can pass a backward-compatibility check while a field changes from “order created time” to “last payment time.” A Consumer can decode the bytes while a sink rejects the value because its domain rule is stricter than the schema.

The contract stack separates Kafka transport, serialization, registry policy, and business meaning

Assign an owner to each layer. The platform team can own the Kafka endpoint, Topic policy, and registry service. Producer teams own publish validity; Consumer teams own the read model and their response to unsafe values. A migration is safer when those responsibilities stay visible.

2Why AutoMQ stays neutral to your serialization

AutoMQ’s Kafka compatibility is valuable here because it preserves the protocol boundary while changing the storage implementation underneath it. AutoMQ’s compatibility documentation describes compatibility with Kafka clients and ecosystem components. That is the right place to start when checking whether a workload’s client and integration surface is supported.

The message path stays outside the registry’s control after serialization. The Producer and its serializer create the key and value bytes. The Kafka client sends them to an AutoMQ Broker. The Broker stores and returns the record. The Consumer and its deserializer interpret the bytes, using the registry when the chosen serialization format requires it. AutoMQ does not need to understand an Avro field definition to persist the record, and a registry does not need to know how AutoMQ lays out durable stream data.

That neutrality is deliberate. AutoMQ uses a Shared Storage architecture in which S3Stream replaces Kafka’s local log storage layer. WAL (Write-Ahead Log) storage provides the durable write path, while S3-compatible object storage holds the primary stream data. Those choices affect broker scaling, recovery, and storage operations. They do not turn the broker into an Avro, Protobuf, or JSON Schema validator.

This separation sets a limit on what an AutoMQ claim should mean. Kafka protocol compatibility can reduce broker-side migration work, but it does not migrate Schema Registry subjects, schema IDs, serializer configuration, policies, or business tests. The registry remains a separate service to deploy, secure, back up, and validate alongside the data plane.

3Run a standard registry beside the Kafka-compatible stack

A practical deployment has two connected paths. Producers and Consumers reach both the AutoMQ bootstrap endpoint and the registry endpoint. The registry reaches the Kafka-compatible cluster when its implementation stores schema state in Kafka, commonly through an internal topic. These paths have different identities, ports, ACLs, and failure modes, so draw them separately in the deployment plan.

For example, a registry service can be configured with the AutoMQ bootstrap servers as its Kafka store and exposed through an internal HTTP endpoint for serializers and deserializers. The exact environment variable names depend on the registry implementation. The important values are the same: the AutoMQ bootstrap address, the registry listener, the authentication and TLS settings, and the permissions needed to create or access the registry’s internal state.

yaml
registry: image: <your-approved-schema-registry-image> http_listener: https://registry.internal:8081 kafka_bootstrap_servers: <automq-bootstrap-server> kafka_security_protocol: <your-client-security-protocol> kafka_auth_config: <your-secret-reference>

Treat this as a deployment shape, not a universal configuration file. A concrete implementation such as Confluent Schema Registry documents its REST resources and Kafka store settings in its Schema Registry API documentation. Use the configuration names, image version, identity, and storage topic rules documented by the registry you have selected. Do not infer that an AutoMQ cluster includes those resources.

The first verification pass should prove the control path before any application rollout:

  1. Confirm that the registry can connect to the AutoMQ bootstrap endpoint and create or read its internal state under the intended ACLs.
  2. Register an initial schema for one representative subject and read it back by subject and version.
  3. Run a compatibility check for an additive change and for a deliberately incompatible change. The registry must return different results for the two cases.
  4. Start a Producer and Consumer with the same registry URL, serializer settings, credentials, and TLS trust material that production will use.
  5. Produce a record, consume it, and record the Topic, partition, offset, schema identity, and application result as one test artifact.

A separate Schema Registry connects to the AutoMQ Kafka endpoint while clients use both paths

This order catches a common mistake: validating a registry with its own API and declaring the platform integration complete. The registry can be healthy while the application uses a different subject naming strategy or the AutoMQ ACL blocks its internal Topic. The end-to-end test has to include the actual client libraries and network path.

4Prove compatibility with real records

Compatibility policy is a statement about reader and writer behavior. A deployment test should therefore use records, not only schema text. Start with a small fixture such as an orders value schema containing an order ID and total. Register version 1, produce records with the production serializer, and save the resulting bytes or a replayable test Topic. Then introduce version 2 in the same way a real release would.

The registry check and the Kafka read test answer different questions. A compatibility endpoint can tell you whether the proposed schema passes the configured rule. It cannot prove that the client is using the expected subject, that the serializer writes the expected envelope, or that a Consumer can read retained records after a restart. Use both checks:

bash
curl -fsS "$REGISTRY_URL/subjects/orders-value/versions/latest" curl -fsS -X POST \ "$REGISTRY_URL/compatibility/subjects/orders-value/versions/latest" \ -H 'Content-Type: application/vnd.schemaregistry.v1+json' \ --data @orders-v2-compatibility-request.json kafka-topics.sh --bootstrap-server "$AUTOMQ_BOOTSTRAP" \ --describe --topic contracts.orders

For Avro, the Apache Avro specification explains writer and reader schema resolution, including defaults and aliases. Test the direction your rollout needs. If Consumers are upgraded first, a proposed reader should handle records written by the previous schema. If Producers are upgraded first, older Consumers must be able to interpret the proposed records, or the release needs a different sequencing rule. Full compatibility protects both directions and usually accepts fewer changes.

The minimum matrix should include these cases:

TestWhat it provesWhy it belongs in the AutoMQ check
Old reader, old recordThe baseline path still worksEstablishes a control result before migration or upgrade
Proposed reader, old recordThe reader can replay retained historySeparates schema resolution from broker storage
Proposed reader, proposed recordThe forward application path worksConfirms serializer, registry, and client settings together
Old reader, proposed recordThe rollout is safe in the other directionRequired when producer and Consumer releases overlap
Replay across offsetsHistorical reads retain expected ordering and visibilityTests the real recovery and backfill path

Run the matrix through the same AutoMQ endpoint and the same registry service that the workload will use. Keep the failure output. A failed subject lookup points to the registry path or client configuration. A failed decode with a successful lookup points to the serializer, schema, or wire format. A successful decode followed by a business rejection belongs to the Consumer’s semantic contract.

Compatibility verification moves from registry policy to producer bytes and consumer reads

The storage architecture still matters to the test. AutoMQ’s S3Stream overview describes separate WAL storage, S3-compatible object storage, and data caching for write and read paths. Test a cold Consumer start, a Catch-up Read over retained records, a registry lookup after local caches are empty, and a Broker replacement or restart within the recovery procedure you intend to operate. The registry check remains valid, but the read path still needs evidence.

5Tombstones and read semantics still hold

A tombstone is a Kafka record with a key and a null value. It is a deletion marker for a compacted Topic, not a schema version containing an empty object. That distinction matters for data contracts because a Consumer must be able to tell “delete this key” from “the value has fields whose contents are empty.” A registry may not perform a value-schema lookup for a null payload, depending on the serializer, so test the chosen client behavior instead of treating tombstones as ordinary value records.

The key must survive. The Consumer must receive the record at its offset, identify the null value, and apply the delete rule for its local state. The Kafka compaction documentation explains the broker’s log-compaction model: compaction runs in the background, so a tombstone does not erase the older record at the instant it is written. A replay can therefore observe an older value and then its tombstone, while a later compacted read may see only the surviving state needed by the application.

That behavior is independent of Avro or Protobuf compatibility. The registry governs schema identity and evolution; Kafka governs the key, null value, offset, partition ordering, retention, and compaction. These semantics live at the Kafka record boundary, so the test should cover them together rather than treating schema validation as proof of correct read behavior.

Use a compacted test Topic and verify four observations:

  • A record with key order-17 and a non-null value can be produced and consumed with the selected schema.
  • A record with key order-17 and a null value reaches the Consumer as a tombstone.
  • The Consumer applies the delete action according to its processing guarantee and commits the expected offset.
  • A replay and a later compaction produce the read behavior documented for that Topic, including the retention window for tombstones.

Do the same for an ordinary retained Topic when the application uses null values as data. The word “tombstone” describes a convention that is meaningful in a compacted log and to a Consumer that understands it. A generic null payload is not automatically a business deletion. The contract must say what the key, null value, offset, and downstream action mean.

6What changes for operators moving from self-managed Kafka

The registry operating model stays familiar across a self-managed Kafka cluster and AutoMQ. You still need subject ownership, compatibility policy, registry backups, client credentials, serializer versions, ACLs, Topic configuration, Consumer lag, replay tests, and a rollback path. AutoMQ changes the storage and broker operating model around that contract, so the runbook should show both the preserved responsibilities and the work that moves elsewhere.

AreaSelf-managed Kafka focusAutoMQ focus
Durable log storageSize broker-local disks, plan replication, and watch partition placement.Validate S3-compatible object storage, WAL storage, cache behavior, and the selected deployment model.
Broker scalingAdd capacity with broker sizing, partition reassignment, and data movement.Size Kafka compute separately from shared storage, then test reassignment, cache warm-up, and traffic distribution.
Failure recoveryRebuild or move local replicas and confirm ISR health.Test shared-storage recovery, WAL recovery, Broker replacement, object storage access, and KRaft metadata behavior.
Registry placementPlace the registry near the Kafka cluster and operate its Kafka-backed state.Keep the registry separate while validating its network, identity, and Kafka store connection to the AutoMQ endpoint.
Schema governanceOperate subjects, compatibility modes, serializer clients, and backups.Keep the same contract controls. Kafka compatibility does not transfer subjects, schema IDs, or policy state for you.
Retention and compactionTie replay windows to local storage capacity and compaction work.Test Tailing Read and Catch-up Read behavior, retention, compaction, cache pressure, and replay cost against the shared storage path.
ObservabilityCorrelate broker disk, replication, client, registry, and Consumer metrics.Correlate Broker, WAL, S3-compatible object storage, cache, registry, client, and Consumer signals across the same test run.

The operator’s center of gravity changes. A self-managed Kafka team protects broker-local capacity and moves partition data safely. An AutoMQ team focuses more on shared storage access, WAL choice, cache behavior, object storage permissions, and the separation between compute scaling and durable history. Both models still require a registry runbook and a compatibility matrix for production clients and serializers.

Before switching a real workload, write down the checks that must stay green:

  • The AutoMQ endpoint accepts the client versions, security settings, Topic configurations, and APIs the workload uses.
  • The separate registry can read and write state through the Kafka path and serve client requests over HTTP.
  • The subject naming strategy, schema IDs, references, and compatibility policy match the production contract.
  • Producers and Consumers pass the version matrix, including defaults, aliases, null values, and tombstones.
  • The team can restore registry state and still decode retained records from the target Topics.
  • Operations can distinguish a registry failure, a serializer failure, a Kafka data-plane failure, and a downstream semantic rejection.

Return to the opening question: protocol compatibility is necessary because it keeps the Kafka client boundary usable, but it is only one layer of a streaming data contract. AutoMQ can preserve that boundary while replacing the storage layer beneath it. The confidence comes from the test matrix around the boundary: a separate registry, real encoded records, explicit tombstones, replayed offsets, and an operating model that names what changes.

To run that matrix against a Kafka-compatible cloud-native streaming platform, start an AutoMQ evaluation with one representative Topic, its registry subjects, and its oldest production reader. The useful result is evidence that the contract survives the platform boundary.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.