Table of Contents
Table of Contents
A schema change can pass review, register successfully, and still break a consumer that is running one deployment behind. The failure often starts with a small ambiguity: when an engineer says “backward compatible,” backward from whose point of view? The producer writing the proposed record, or the consumer reading records already in Kafka?
That question has a precise answer. Backward and forward compatibility describe the direction in which a reader and a writer must continue to work across schema versions. The useful policy is the one that matches the rollout order, the retained data window, and the serializers that actually run in production. This article uses four anonymous, reproducible illustrative scenarios. They describe common failure patterns, not named customers, reported incidents, or customer data.
1Backward and forward compatibility protect opposite readers
Start with two schemas and two readers. S_old is the schema that has already written records, and S_proposed is the schema being introduced. A reader resolves a writer's bytes through its own reader schema. The direction is named after the reader that must survive the change:
| Direction | Reader and writer relationship | Rollout it protects |
|---|---|---|
| Backward | A reader using S_proposed can read records written with S_old. | Upgrade consumers before producers. |
| Forward | A reader using S_old can read records written with S_proposed. | Upgrade producers before consumers. |
| Full | Both relationships work. | Allow a mixed window in either rollout order. |
| None | The registry does not enforce a compatibility relationship. | An explicit migration or application test owns the risk. |
The word “backward” therefore points from the proposed reader back to the records already present. “Forward” points from an existing reader forward to records written by the proposed producer. A registry may call these modes BACKWARD, FORWARD, and FULL, but the names do not replace a rollout decision.
Consider a consumer group that must be upgraded before its producer. The group will see old records after the reader deployment and before the producer changes. Backward compatibility is the relevant gate. Reverse the order and forward compatibility becomes the gate. If producers and consumers can be deployed independently, full compatibility buys a wider safety margin, at the cost of rejecting more schema changes.
The distinction also applies to replay. A consumer may read records written several schema versions ago because Kafka retention has kept them, a sink is rebuilding a table, or an incident response run is replaying a partition. A check against only the immediately prior version may be too narrow for that read path. Use the registry’s transitive option, or an explicit batch that tests every writer version the consumer can encounter.
Full compatibility is a useful constraint, but it is not a permanent approval for every future change. It says that the tested reader-writer relationships work under the selected serialization rules. It does not approve a changed unit, a newly overloaded enum value, a different subject naming rule, or a consumer that silently treats an unknown value as a known one. It also does not keep an old serializer configuration available after a client upgrade. Keep the mode, the rollout order, and the semantic assumptions together in the change record. A mode name tells the registry which structural questions to ask; the application test tells the team whether the answer is safe for its actual read path.
2Read and write rules that survive a mixed deployment
A compatibility policy becomes useful when it is translated into two concrete promises. The writer promises what it will emit during the coexistence window. The reader promises which prior and proposed records it can resolve and what it will do with missing, null, defaulted, or unknown values.
For an Avro record, the Apache Avro specification defines the reader-writer resolution rules. A reader can supply a default when an older writer omitted a field. An older reader can ignore a field that a proposed writer added, provided the serializer and schema rules for the event format support that behavior. Those structural rules still need a business decision: a default of "USD" is safe only when the historical records really mean USD, rather than “currency was not captured.”
A practical contract for a mixed deployment looks like this:
- Add a field with a valid reader default before any producer depends on it. For a nullable Avro field, use a union whose first branch matches the default, commonly
null, and document what missing means. Treat “missing” and “unknown” as deliberate values when the domain needs that distinction. - Keep a field while old readers, sinks, replay jobs, or rollback paths can still encounter it. Removing a producer field is a forward-compatibility decision, even when the proposed reader no longer uses it.
- Use an alias for a rename only when the business meaning is unchanged. If the name changes because the field now identifies a different entity, create a visible event or field transition instead of hiding it behind an alias.
- Test the actual serializer, deserializer, subject naming strategy, wire envelope, and registry policy. A schema file that parses in source control does not prove that a running client can decode a Kafka record.
- Choose latest-only or transitive checking from the replay window. A retained record is part of the contract whether or not it is part of the latest deployment.
These rules separate structure from meaning. Schema Registry can reject an incompatible field shape, but it cannot tell whether a default is truthful or whether account_id and customer_id identify the same thing. That semantic check belongs in the contract review and the consumer fixtures.
3Four illustrative scenarios that fail in reusable ways
Each scenario below can be reproduced with two schema fixtures, an actual serializer, and a reader-writer matrix. The team names are descriptive labels only. They do not identify real companies.
3.1Scenario 1: “The reader was safe, but the required field was not”
An illustrative payments team adds currency as a required string to an Avro Payment record. Its consumer deployment goes first because the team wants to read the field before producers start sending it. The consumer can read the proposed schema, but it cannot resolve the older records that contain amount_cents and no currency. The first replay exposes the problem, even though the producer has not emitted a bad record.
The mistake is a backward-compatibility failure. The proposed reader is the party that must read old data, and it has no rule for the missing field. The repair is to add a nullable union with a valid default, commonly null, or another field type with a valid reader default, then run the proposed reader against representative old records before deployment. If the business cannot define a truthful default, the team should keep the field absent and make the event change visible through a versioned contract.
3.2Scenario 2: “The producer moved first and old readers met a missing field”
An illustrative logistics team removes delivery_window from the writer because its replacement field, promised_at, is ready. The producer rolls out first. An older warehouse consumer still expects delivery_window as a required value, so deserialization fails as soon as it receives a proposed record. The topic remains healthy; the reader is the component that breaks.
This is a forward-compatibility failure. The old reader must understand the proposed writer’s output. The fix is to keep delivery_window during the mixed window, emit a value that preserves the old contract, or move the meaning change to a separate event version. The retirement condition should be observable: all known readers, sinks, and replay jobs have moved, and the rollback plan no longer needs the old field. A successful registration check under backward mode would not have protected this producer-first rollout.
3.3Scenario 3: “The alias was missing, or the meaning had changed”
An illustrative account team renames account_id to customer_id to make the domain language clearer. An old consumer still reads account_id, and the proposed producer emits only customer_id. The two fields are both strings, so a type check looks harmless. The reader still cannot find the old field name in the proposed record.
If the identity is unchanged, the team can deploy a reader that accepts the previous name through an Avro alias, keep that alias through the retained-data and rollback windows, and change the producer only after old readers, sinks, replay jobs, and rollback paths have moved or are otherwise protected. If the identity really changed, an alias would conceal a semantic break. In that case, publish a separate field or event version and make consumers choose the transition explicitly. Compatibility checks can resolve names; they cannot approve a changed business definition.
3.4Scenario 4: “Latest passed while replay still failed”
An illustrative analytics team has three retained writer versions. A proposed schema is checked against the latest version and passes. A replay job then reads an older record whose field was represented differently, or whose referenced record definition is no longer available under the expected subject. The live consumer had no error because it had seen only the latest shape. The replay path fails later, under a different reader configuration.
The failure is a test-scope problem that becomes a compatibility incident. The repair is to make the historical window explicit, use transitive checks where the registry supports them, export referenced schemas with their subject mappings, and run a batch decode against actual encoded fixtures from every retained writer version. Include sink and backfill consumers in that matrix. “Compatible with the latest version” is a narrow statement; “compatible with the records this job can read” is the operational contract.
The four patterns are different, but the diagnostic sequence is stable:
- Identify the writer schema that produced the bytes.
- Identify the reader schema and deployment order.
- Name the field, type, alias, default, or subject mapping that must resolve.
- Reproduce the failure with the real serializer and a record fixture.
- Repair the rollout order or contract, then keep the fixture as a regression test.
4Run compatibility as a batch, then enforce it in CI
A registry check is one layer of the test. A compatibility batch should exercise the path the application will use, including the wire format and the records that already exist. A small matrix is more informative than a large collection of latest-schema snapshots:
| Test pair | What it proves | Typical failure |
|---|---|---|
| Old writer → old reader | The baseline fixture still works. | Bad fixture or serializer drift. |
| Old writer → proposed reader | Backward compatibility. | Missing default, alias, or reference. |
| Proposed writer → old reader | Forward compatibility. | Removed field, renamed field, or unsupported type change. |
| Proposed writer → proposed reader | The intended rollout path works. | Serializer, subject, or wire-format mismatch. |
| Every retained writer → current reader | Replay safety. | Latest-only check hid an older break. |
Run this batch when a schema changes, a serializer library changes, the registry policy changes, or a Kafka-compatible platform changes. Store the encoded fixtures with the schema text and the subject/version metadata. That makes a failure inspectable rather than dependent on a running registry retaining the exact historical state.
CI should reject a change before application deployment, but it should not pretend that structural compatibility is the entire contract. A useful pipeline has separate checks for the registry policy, serializer round trips, semantic fixtures, and the deployment matrix:
pull request
-> compare proposed schema with latest and transitive policy
-> encode representative old and proposed records
-> decode with old and proposed readers
-> assert defaults, null rules, aliases, and unknown-value behavior
-> run retained-record replay fixtures
-> publish the compatibility report with the change
The report should state which direction was required and why. “Backward passed” is incomplete without “consumers deploy before producers” or the reverse. It should also name the oldest writer version covered, the serializer and registry configuration used, and any scenario that requires a separate semantic approval.
At deployment time, repeat a smaller smoke batch against the target environment. Confirm the subject exists, the client can resolve its schema, a proposed record can be read by the intended consumer, and a retained record can still be decoded. For a Kafka-compatible target, run the same client and serialization matrix against the target Kafka endpoint. The schema registry remains a separate governance component, while the data plane must still carry the records and client behavior your application depends on.
This boundary is reusable with AutoMQ. AutoMQ is Kafka-compatible, so teams can evaluate their existing Kafka clients, serializers, and standard Schema Registry setup against its Kafka endpoint. AutoMQ does not require a proprietary schema feature for the compatibility model described here. Its Apache Kafka compatibility documentation describes the protocol and ecosystem boundary; its architecture overview describes the Shared Storage architecture underneath that boundary. The schema policy, serializer fixtures, and rollout decision still belong to the application team.
For adjacent ownership guidance, see data contracts in practice and schema registry runbook guardrails. Both are useful after the direction test has exposed who owns the contract and the operational response.
5A quick reference for the next schema change
The shortest reliable decision is to name the reader that must survive the rollout. If consumers move first, protect the proposed reader with backward compatibility. If producers move first, protect the existing reader with forward compatibility. If both orders are possible, test full compatibility or constrain the deployment sequence.
| Change | Reader risk | Safer sequence | Gate to run |
|---|---|---|---|
| Add a field | Proposed reader sees it missing in old records. | Add a valid default or explicit nullable meaning, deploy readers, then writers. | Old writer → proposed reader. |
| Remove a field | Old reader expects it in proposed records. | Stop depending on it, keep it during the mixed window, then retire it. | Proposed writer → old reader. |
| Rename a field | Reader cannot map the old name. | Use an alias when meaning is unchanged; version the event when meaning changes. | Old and proposed readers against both writer forms. |
| Change a type or reference | A serializer or replay path cannot resolve the representation. | Test the actual wire format and every retained writer version. | Full writer-reader matrix plus replay. |
Return to the original question: backward from which side? The answer is the reader that must keep working. Write that sentence in the change record, attach the old and proposed fixtures, and make CI prove the direction before a deployment turns a schema decision into a consumer outage. When you want to run the same compatibility batch on a Kafka-compatible streaming platform, start an AutoMQ evaluation.
