Blog

Breaking Change Week: How Schema Evolution Goes Wrong Without a Compatibility Policy

Table of Contents

Table of Contents

The incident starts with a harmless looking pull request: remove a field that has been empty for months. The producer deploys, the topic accepts records, and the first consumer keeps processing. Then a replay job starts reading older events. A connector sees a shape it did not expect. A dashboard goes quiet because its transformation treated the missing field as a failure. By Friday afternoon, the team is debating whether the schema, the serializer, the consumer, or the rollout was at fault.

The root problem is usually earlier and less dramatic. Nobody wrote down what “compatible” means for the stream, which versions may coexist, who can approve an exception, or how long an old field must remain readable. A registry can reject some incompatible shapes, but it cannot supply the rollout order or the business meaning of a field. A Kafka schema compatibility policy closes that gap by turning schema evolution into an operational contract.

The useful policy is small. It names the default, defines the exceptions, blocks unsafe changes in CI, and gives old readers a deliberate retirement window. The same contract can protect a Kafka deployment on conventional brokers or a Kafka-compatible platform such as AutoMQ, because the risk sits at the producer, serializer, registry, and consumer boundaries.

Breaking change timeline showing an unchecked schema commit reaching production and failing during replay

1The breaking change that started on a Friday

Consider an order event with order_id, status, and total. Several producers write it, a fraud service reads it in real time, a warehouse connector loads it, and a support tool replays it when an agent opens a case. The field status_reason looks like a safe removal because the primary consumer stopped using it. The producer team deletes it, passes unit tests, and ships.

What the team did not test was the whole lifetime of the event. The warehouse mapping still expects the field for a historical backfill. The support tool uses a missing value differently from an empty value. A second consumer has an older reader and cannot resolve the new writer schema. The stream never stopped accepting records, so the broker offered no signal that the contract had changed underneath its consumers.

There are two separate failures here. The structural change may violate the selected schema compatibility mode. The rollout also violated an application contract by removing a field before every reader, replay path, and rollback path had stopped depending on it. A policy must address both. Otherwise a team can fix the registry setting and still ship a semantically breaking event.

A configured registry is not governed evolution. It checks versions; policy covers rollout order, retention, replay, and ownership.

2Evolution traps that appear without a policy

Teams without a written policy tend to repeat the same mistakes because each producer team solves a local problem. The following changes are not automatically unsafe, but each needs an explicit compatibility direction and a rollout plan.

2.1Removing or renaming a field in one release

A field removal can be structurally compatible for one reader and operationally unsafe for another. A rename is even more subtle: a reader may need an alias or a translation layer to interpret old records, while old writers continue to emit the previous name. Treat a rename as an add-and-drain sequence. Introduce the new field, let readers accept both names, migrate writers, observe the old name, and remove it after the replay and rollback conditions are satisfied.

2.2Changing meaning while keeping the same type

Changing total from cents to dollars may leave the serialized type untouched. A registry will not infer that the unit changed. The same problem appears when status = "pending" changes from “awaiting payment” to “awaiting manual review.” Structural compatibility cannot protect semantics that are not documented. A policy therefore needs a semantic review for unit, enum, nullability, key meaning, and timestamp interpretation changes.

2.3Making an existing field mandatory

An older producer will not start populating a field because a newer reader requires it. If the reader cannot supply a safe default, the rollout order is wrong. Add the field as optional or with a business-valid default, deploy readers that tolerate both shapes, and then make the producer write it. A syntactically valid default is not automatically a truthful default. “Unknown” is often safer than inventing a value that looks real.

2.4Updating the registry without updating the wire path

Schema files, subject names, serializer settings, registry URLs, and wire-format expectations travel together. During a migration, a producer can still connect to Kafka while its serializer points to the wrong registry or subject. A consumer can fetch records while failing to resolve the schema ID. The compatibility check must exercise the actual serialization and deserialization path used in deployment, including authentication and subject naming.

The schema diff is one input to a release decision. The policy must assign semantic ownership, define registry enforcement, state what CI proves, and name the observation that permits the next step.

3A minimal Kafka schema compatibility policy

The following text is deliberately short enough to put in a repository next to the schema definitions. Adapt the names and time windows to your retention and rollback requirements, but keep the decisions visible.

Default. Every production subject uses a declared compatibility mode. The default is backward compatibility when consumers are upgraded before producers. A team choosing forward or full compatibility records the rollout order and the reason.

Required checks. A schema change must pass registry compatibility validation, serialization and deserialization fixtures for every supported reader version, and semantic review for units, enum values, nullability, keys, and field meaning.

Allowed evolution. Additive fields require a reader-safe default or an explicitly nullable representation. Renames use an alias or a translation period. Removals and type changes require an approved migration plan. Changing business meaning under the same field name is a breaking change even when the wire type is unchanged.

Release gate. CI rejects a change that fails compatibility, lacks fixtures for retained versions, or has no named owner and rollback path. An exception requires a written impact list, a consumer communication record, and an approver outside the producer team.

Retirement. Old fields and reader behavior remain supported until the removal window ends. The owner must show that supported consumers, replay jobs, connectors, and rollback versions no longer require the old shape before removing it.

Minimal schema compatibility policy card covering defaults, checks, rollout, and retirement evidence

This policy gives teams a default and makes deviations deliberate. It also separates structural checks from semantic review. No compatibility mode can decide whether a default is honest or an enum value changes a business decision.

Choose the compatibility direction from the deployment sequence, not from habit. Backward compatibility protects a new reader consuming old records. Forward compatibility protects an old reader consuming records written with the proposed schema. Full compatibility asks for both. If consumers can replay records from more than the latest version, use a transitive check or an equivalent fixture set that covers the retained history. The mode is a release constraint, not a permanent substitute for ownership.

4CI gates that catch the break before deploy

The policy becomes real when a pull request can fail before a producer publishes the first new record. Keep the gate close to the schema change, but make it use the same subject, registry configuration, serializer, and generated artifacts as the application build.

The schema compatibility gates guide covers the wider production checklist around this boundary. This article narrows that work to the policy text and rollout evidence that a producer team can attach to a change.

The minimum gate has four inputs:

  1. The proposed schema and its declared subject. This catches accidental subject creation, naming drift, and a compatibility mode that was never recorded.
  2. The previous and supported reader schemas. Comparing against the latest version alone misses a replay path that still reads an older retained event.
  3. Wire-level fixtures. Serialize representative old and proposed records, then deserialize each with the readers that the rollout permits. Include missing fields, defaults, aliases, nulls, and enum values where they matter.
  4. Ownership and rollout metadata. Require a topic owner, consumer impact list, deployment order, rollback version, and removal-window owner. A green schema diff without these fields is an incomplete release.

The gate should report a useful failure. “Compatibility check failed” forces the producer team to investigate the tool. “Reader warehouse-v4 cannot resolve status_reason from the retained fixture; either keep the field or attach an approved migration” points to the decision. Store the proposed schema, check result, fixture versions, and approval with the build so an incident responder can reconstruct what was allowed.

The CI result is a safety boundary, not a substitute for production observation. It cannot prove that an unregistered JSON consumer will interpret a value correctly, or that a connector will preserve a renamed field. For high-impact subjects, add a contract test against a production-shaped environment and make the test publish and consume through the same registry path as the service.

CI gate flow showing schema diff, compatibility fixtures, semantic review, and a blocked deploy

5Gray reads, dual reads, and the removal window

A compatible schema still creates a mixed-version period. Treat that period as part of the design. The safest rollout usually gives readers more tolerance before writers start using the new shape.

For an additive field, deploy a reader that can handle the old record first. It should distinguish “field absent” from a meaningful business value when that distinction matters. Then update writers, observe the new field in the consumer path, and keep the reader able to process both forms. This is a gray read: the reader knows the old and proposed shapes are expected during the transition. The broader schema registry operating model helps place ownership and audit work around the same rollout.

For a rename, dual-read behavior is often clearer than a sudden replacement. The reader checks the new field and falls back to the old field while both exist. If both are present, it applies a documented precedence rule and emits a metric or structured log for disagreement. Writers can then populate both fields, move consumers, and stop writing the old name after the acceptance criteria pass. The old reader path remains available for rollback until its owner signs off.

For a type or semantic change, use an explicit versioned field or event envelope when one reader cannot safely interpret both meanings. A total that changes unit should not be hidden behind a compatibility setting. Keep the old meaning available, write the new meaning under a distinct contract, and let consumers migrate against fixtures that show the conversion. The extra field is safer than silently corrupting a downstream calculation.

The removal window is the part teams skip when the rollout looks healthy. Define its end with evidence:

  • All declared consumers have deployed a reader that accepts the proposed shape.
  • Replay jobs and connectors have processed representative old and new fixtures.
  • The producer no longer writes the old field, or the old field has a documented translation path.
  • The rollback version can still read records written during the mixed window.
  • The owner has checked retention, scheduled jobs, and dormant consumers that do not appear in the live dependency graph.

Do not use elapsed time alone as proof. A week with no traffic from a batch consumer says little if that consumer runs monthly. Tie removal to the longest supported replay and rollback path, then record the evidence with the change. If the team cannot identify that boundary, the field is not ready to be removed.

6The AutoMQ boundary: keep the policy above the broker

The compatibility policy should travel with the Kafka workload. AutoMQ is a Kafka-compatible cloud-native streaming platform, so producers, consumers, serializers, and Schema Registry checks remain part of the application contract. Its Kafka compatibility documentation is the place to verify the supported Kafka surface for a specific workload.

That boundary matters during a platform review. Schema evolution does not become safe because the broker uses local disks, tiered storage, or shared storage. The producer and consumer still need the same policy, fixtures, rollout order, and removal evidence. AutoMQ's shared-storage architecture changes how durable stream data is stored and how broker lifecycle is separated from storage, while the schema gate remains an application and registry concern.

This separation is useful when the same contract must survive a platform change. Keep the subject names, schema artifacts, fixtures, reader behavior, and CI results under version control. Test the complete client and serializer path against the target Kafka-compatible environment. A platform migration can preserve Kafka-facing application behavior, but it does not excuse missing registry configuration, hidden consumers, or an untested replay path.

If a team is evaluating AutoMQ for an existing Kafka workload, start with one subject that has real retention, a real replay job, and a known rollback path. Run the compatibility gate against old and proposed records, then measure whether the deployment boundary changes any client or registry assumptions. The result should be evidence for a policy decision, not a claim that the platform removes schema governance.

7A rollout checklist that survives the next Friday

Put the following in the pull request template for every production schema change:

  • Name the subject, topic owner, supported writer versions, and supported reader versions.
  • State whether the change is additive, a rename, a removal, a type change, or a semantic change.
  • Declare the compatibility mode and rollout order.
  • Attach registry validation and wire-level fixture results.
  • List live consumers, replay jobs, connectors, and rollback readers.
  • State how gray reads or dual reads behave when both fields are present or absent.
  • Define the removal evidence and name the person who can approve cleanup.
  • Link the build artifact, approval, and post-deploy observation to the contract record.

The team that owns a topic does not need to predict every future consumer. It does need to make the current contract explicit and give future readers a safe way to enter. That is the difference between a schema file and a compatibility policy.

When the next “small” field change arrives, the goal is a boring CI failure if the change is unsafe, a quiet mixed-version rollout if it is safe, and a removal decision backed by evidence. Put the policy beside the schemas, run it through the same registry and serializer path used in production, and test it on the Kafka-compatible platform that carries the workload. You can also explore AutoMQ with a workload-specific review, bringing the same compatibility policy, fixtures, and rollback questions to the evaluation.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.