Table of Contents
Table of Contents
There is a schema in the repository that nobody wants to touch. It is the one that passed review, reached production, and then turned out to describe an event incorrectly. The field is technically present, the serializer still produces bytes, and the registry still returns a valid schema ID. The failure appears later, when a consumer interprets a value with the wrong meaning or a replay job reaches records written during the bad deployment.
The first instinct is often to remove the bad schema version. That feels tidy. It is also where a recoverable application mistake can become a data recovery problem. Deleting a registry entry does not rewrite records already stored in a Kafka topic. It can make those records harder to decode during a replay, backfill, audit, dead-letter investigation, or consumer restart.
A safer production response treats a bad schema as something to neutralize. Stop it from creating more damage, preserve enough history to decode what already exists, and move readers and writers through an explicit transition. The practical choices are a compatible correction, a tombstone or hidden-version policy, and a dual-read window. Each has a different precondition.
1The schema you cannot delete
A Schema Registry usually manages a subject’s schema versions and assigns identifiers that serializers place into encoded records. A version is registry history. It is not a Kafka offset, a Topic, a Partition, or a pointer to a range of records. The registry can tell a consumer which writer schema an ID refers to; Kafka stores the encoded record and its headers as opaque data.
That boundary explains the operational trap. Suppose a producer serialized records with a bad version before the team noticed the error. Removing that version from the registry does not remove those records from the topic. It only changes what a serializer or deserializer can discover when it needs the writer schema again. A warm client may continue to work from its local cache while a cold consumer, a replay tool, or a different language client fails on the same historical record.
The reverse mistake is also common: treating a Kafka tombstone as if it deletes a schema version. A Kafka tombstone is a record with a key and a null value, used with log compaction to express that the current value for that key should be removed from the compacted result. It belongs to Topic data. It does not delete, hide, or invalidate a Schema Registry version. Kafka’s log compaction documentation describes that record-level behavior.
A registry version can be marked unusable by an operating policy, a deployment control, or an implementation-specific delete or deprecation workflow. The word “tombstone” is useful as shorthand for that policy only if the team defines it clearly. It is not a universal Schema Registry API operation, and it should never be used to imply that an existing version has vanished.
A rollback should change who may use a schema next. It should preserve the information needed to understand records already written.
Before choosing a pattern, freeze the facts:
- Which subject and schema identifier are involved?
- Did the producer emit records with the bad version, or did validation stop the rollout first?
- Which consumers, connectors, batch jobs, and replay tools still read the subject?
- Can the registry retain the version while preventing fresh writes?
- Is the defect structural, such as an incompatible field change, or semantic, such as a valid field whose meaning is wrong?
- What is the longest replay or retention obligation for the affected Topic?
That inventory matters because “rollback” has two separate targets. You may need to stop future serialization, or you may need to make old records readable again. A registry action can solve the first and leave the second untouched.
2Why deletion is the dangerous option
Deletion is attractive when the bad schema has not been used. If no record carries its identifier and every producer can be verified against an earlier version, removal may be a reasonable cleanup step in a registry that supports it. Production evidence, rather than the existence of an embarrassing version, should make that decision.
Once records exist, deletion has a wider blast radius:
- Replay loses its lookup path. A consumer reading from an older offset may need the writer schema to resolve fields or defaults. If the registry can no longer return it, the consumer may fail even though the Kafka record is intact.
- Caches hide inconsistent states. One process may have the schema locally; another may need a registry lookup. The incident can look intermittent until a restart or rebalance clears the cache.
- Recovery tooling is often colder than production. Backfills, export jobs, and forensic scripts may not run often enough to hold every schema in memory. They are the first clients to reveal that registry history was removed.
- Removing metadata does not reverse meaning. If a field named status was populated with the wrong business meaning, deleting the version cannot repair the values already serialized. It only makes the historical contract less visible.
Soft and permanent deletion can have different behavior by implementation and configuration. Verify the registry API before using either operation, and keep both separate from deleting Kafka data or planning for records already in the log.
The safer default is therefore:
- Preserve the schema and its identifier while the incident is being scoped.
- Block the producer path that can create more records with the bad version.
- Decide whether old records must remain readable, be quarantined, or be transformed.
- Record the retirement rule separately from registry storage, so “do not use” does not become “cannot decode.”
Historical clutter is a smaller operational cost than discovering after a consumer restart that the last copy of a writer schema was removed.
3Pattern one: compatible correction
A compatible correction works when the defect can be fixed by registering a later schema that preserves the contract expected by existing readers or writers. The registry’s compatibility mode is the guardrail, but the mode does not prove business meaning. A field can keep the same type and still carry the wrong unit, code, or interpretation.
The correction path looks like this:
- Stop or gate producers that would continue using the bad version.
- Capture the exact schema text, subject, identifier, and producer build that emitted it.
- Define the corrected schema and test it against the compatibility direction used by the subject.
- Register the corrected version under the intended subject.
- Roll producers forward to the corrected version.
- Keep readers able to handle the historical version until the affected records leave the replay window.
Compatibility is strongest when the problem is structural and the corrected shape can satisfy the configured reader or writer contract. It is weaker when the defect is semantic. Adding a field with a safe default may pass a backward-compatibility check, while changing a value from cents to dollars will not be caught by a type comparison. That case needs a semantic guard in the application, an added field or subject, or a transformation step.
Preconditions
- The registry compatibility mode and subject naming strategy are known.
- The corrected schema passes tests with representative legacy and corrected records.
- The serializer can be pinned to the corrected identifier instead of reusing a cached bad mapping.
- Consumers have a plan for records already written with the bad version.
- The team can monitor both registration and deserialization errors during rollout.
This pattern is the least disruptive when no semantic ambiguity remains. It does not make the bad version disappear. It makes the safe path the active path while retaining a decode path for history.
4Pattern two: tombstone the use of a version
A tombstone or hidden-version pattern is a policy state, not a deletion of registry history. The team keeps the schema version available for decoding and marks it as forbidden for fresh production writes. Admission controls can reject that identifier, block the producer release that requests it, or route the subject to an approved version. The exact enforcement point depends on the serializer, registry, and deployment system.
The key is to separate two questions:
- May a producer create another record with this version?
- May a reader resolve this version while processing an existing record?
The first answer can be no while the second remains yes. That is the useful asymmetry. It lets operators stop the leak without burning the map.
A version-hidden policy should include an owner, a reason, and an exit condition. A deployment record can block production writers because a field used the wrong semantic code while readers retain support until affected offsets are replayed or transformed. The policy belongs in the operating workflow even if the registry has no native “tombstone” state.
Preconditions
- The team can identify the bad version at the producer boundary.
- Producers and serializers expose enough information to reject or route that version.
- Readers can still resolve the retained schema ID.
- There is an audit trail for the block, its owner, and its removal condition.
- The organization has a replay or retention statement for the affected Topic.
This pattern is useful when the schema is valid enough to decode but unsafe to emit. It is also the cleanest choice when the team needs time to coordinate consumers. It does not repair values that are already wrong. If the version encoded invalid business meaning, pair the block with a correction, a quarantine path, or a data transformation plan.
5Pattern three: dual-read during the rollback window
Dual-read is the right shape when bad records already exist and the system needs to move from one contract to another without stopping every consumer at once. Writers switch to the corrected schema, while readers accept both the historical and corrected forms for a bounded period. The reader can branch on schema ID, normalize both forms into one internal model, and emit a metric for every record decoded through the legacy path.
This is a migration window, not a permanent invitation to support every schema ever published. Set a start condition, an end condition, and an owner. The end condition might be a verified offset boundary, completion of a backfill, expiration of the relevant retention policy, or confirmation that no approved replay tool still needs the legacy form. Avoid using wall-clock age alone when consumers can reset offsets or when a compacted Topic retains keys beyond the normal processing path.
A dual-read rollout usually needs four controls:
- Write selection: subsequent records use the corrected schema after the producer cutover.
- Read branching: consumers select a decoder by schema ID or a compatible reader strategy.
- Semantic normalization: both forms map to the same internal field meaning, with explicit conversion where needed.
- Exit monitoring: legacy decode count, deserialization failures, replay progress, and consumer lag determine whether the window can close.
Preconditions
- The consumer can inspect the schema identifier or otherwise distinguish the two encodings.
- The application team can write and test two decoder paths.
- The normalized model has an explicit rule for fields whose meaning changed.
- Producers can cut over independently from consumers.
- The team can replay representative records before and after the switch.
Dual-read does add code and test surface. That cost is justified when the alternative is either dropping historical data or forcing a synchronized stop across unrelated consumers. Once the legacy count reaches zero under the agreed replay boundary, remove the old read path in a separate change. Keep the registry history long enough for the rollback procedure itself to remain repeatable.
6What changes when the broker is Kafka-compatible?
These patterns live at the serialization and registry boundary. They do not depend on a proprietary broker-side interpretation of the schema. A Kafka-compatible streaming platform can carry the encoded records, offsets, Topics, and consumer groups while the registry and clients enforce the schema contract.
AutoMQ follows that boundary as a cloud-native streaming platform compatible with the Apache Kafka protocol. Its Kafka compatibility documentation is the relevant place to verify client and ecosystem assumptions. The rollback decision remains the same: use the registry’s standard subject, version, and identifier semantics, and treat Topic records as data that must remain readable.
AutoMQ’s architecture documentation explains its Shared Storage architecture, but that storage design does not turn schema versions into Topic records or give a broker the authority to erase registry history. This is a useful portability test. If a rollback requires undocumented registry behavior, it is coupled to the registry workflow, not to Kafka storage.
7A rollback decision table
Use the narrowest pattern that satisfies the evidence. If the evidence is incomplete, preserve the version and treat it as blocked until the inventory is complete.
| Observed state | Preferred pattern | Preconditions to verify | Do not close the incident until |
|---|---|---|---|
| No records used the bad version and every producer is accounted for | Compatible correction or controlled cleanup | Producer evidence, registry audit, and compatibility tests are complete | A fresh canary record decodes correctly |
| Records exist, but the schema shape is readable and unsafe only for subsequent writes | Tombstone or hidden-version policy | Writer admission can be blocked while reader lookup remains available | Bad-version write attempts are rejected and historical reads still work |
| Records exist and consumers cannot cut over together | Dual-read window | Both decoder paths, semantic normalization, and replay tests are ready | Legacy reads fall below the agreed boundary and errors remain controlled |
| Historical usage is unknown, or replay and audit obligations are open | Preserve and investigate | Subject, IDs, producer builds, consumer groups, and replay tools are inventoried | The owner can prove which readers still need the version |
| The problem is semantic data corruption, not an encoding mismatch | Block plus correction or transformation | Business meaning is defined and affected records are identified | A repair or quarantine path has been tested on representative records |
The table separates two kinds of rollback that teams often collapse into one. Registry rollback chooses which schema can be registered or used going forward. Data rollback decides what to do with records already in Kafka. The first may be a metadata or policy change. The second may require a dual-read adapter, a backfill, a quarantine Topic, or a business decision to preserve the original values.
The repository copy of the bad schema is not a mistake to erase. It is part of the evidence needed to decode, explain, and if possible repair the records it produced. When a production schema goes wrong, preserve the history, block the unsafe path, and make the rollback boundary explicit. If your team wants to test these registry and Kafka compatibility assumptions on a storage architecture built for cloud operations, explore AutoMQ.
