Table of Contents
Table of Contents
A schema change is easy to approve and hard to replay. A producer can register a compatible version, pass its deployment checks, and still leave a consumer unable to decode data during a long backfill. In a diskless Kafka design, the same decision crosses another boundary: the records may outlive the broker that first served them, while the schema registry, cache, metadata, and object-storage paths have different failure and retention rules.
The production question is therefore larger than “is this schema backward compatible?” It is whether every consumer that may read the record can resolve the right schema for as long as the record remains part of the service contract. That contract has four parts:
- Schema policy: who can register a version, which compatibility mode applies, and how the check runs in CI/CD.
- Kafka data: record bytes, headers, offsets, timestamps, and the schema identifier carried by the chosen serializer.
- Storage path: broker cache, write-ahead log (WAL) or buffer, and the object-storage range that may serve a replay.
- Reader lifecycle: current consumers, paused consumers, recovery jobs, connectors, and analytical backfills.
The thesis is practical: diskless Kafka does not change the meaning of a schema version, but it makes the lifetime and serving path of old versions more visible. Treat schema evolution as a replay and recovery contract, then test it across the storage boundaries.
1What schema evolution means in a diskless Kafka design
Schema evolution is the controlled change of a data contract while existing producers and consumers continue to interoperate. A registry stores versions and applies a compatibility rule before accepting a candidate version. For example, Confluent’s Schema Evolution and Compatibility documentation distinguishes backward, forward, full, and transitive checks. The names are useful, but the policy is only meaningful when it matches the readers that can still encounter an older record.
A registry is not the Kafka log. The serializer usually places a schema identifier in the encoded record, while the registry stores the schema definition and its compatibility history. Kafka retains the encoded record according to topic policy. A consumer may therefore need both the old record and the old schema after a broker has evicted its local cache or after a replacement broker begins serving a historical range.
That separation gives a better review question:
When a consumer replays a record from shared storage, which component resolves its schema, and what happens if that component is unavailable or has already removed the version?
Answer it before discussing a diskless implementation. The answer should name the registry endpoint, the client cache behavior, authentication, schema deletion policy, and the fallback or alert when resolution fails.
| Contract element | What changes over time | Evidence to collect |
|---|---|---|
| Schema versions | Fields, defaults, enum values, unions, or references | Registry version history and compatibility decision |
| Record bytes | Serialized payload and schema identifier | Producer serializer configuration and sample payloads |
| Kafka history | Retention, replay range, offsets, and timestamps | Oldest readable offset and backfill scenarios |
| Serving state | Broker cache, WAL or buffer, and object-storage ranges | Cache misses, fetch latency, upload status, and retries |
| Reader population | Current, paused, newly deployed, and recovery consumers | Consumer versions, connector converters, and restore runbooks |
A schema check against only the newest consumer is incomplete. Inventory every reader that can consume from the topic: stream processors, Kafka Connect workers, batch jobs, data quality checks, incident tools, and one-off replay scripts. Record which serializer and registry client each one uses. A consumer that is not part of the deployment pipeline is often the one that fails during a replay.
Compatibility modes also need an explicit direction. Backward compatibility asks whether a updated reader can read old data. Forward compatibility asks whether an old reader can read candidate data. Full compatibility combines both directions, while transitive checks compare against the relevant history instead of only the latest version. The correct setting depends on the rollout order and how long old readers may remain active; it is not a storage-tier property.
2The mechanism: registry, cache, metadata, and object storage
A diskless Kafka data path has two related but independent control loops. The registry governs how a schema version is admitted and resolved. The Kafka storage path governs where encoded records and their indexes are written, cached, uploaded, fetched, and eventually deleted. A green signal in one loop does not prove the other is ready.
The write path can be traced as a sequence:
- A producer serializes a record with a registered schema identifier.
- The Kafka leader accepts the record under the producer’s acknowledgement and durability settings.
- The broker’s configured write path makes the record available for the acknowledgement contract.
- The record and its indexes move into the shared durable layer and become available for later reads.
- A consumer fetches the record, resolves the schema through its registry client, and decodes the payload.
Each step has its own evidence. A successful registration proves that the proposed schema passed the registry rule; it does not prove that a paused consumer can resolve the version. A successful Kafka write proves that the record met the broker’s acknowledgement contract; it does not prove that a backfill worker has the required registry permission. A successful object upload proves durable bytes exist; it does not prove that a cache eviction or cold fetch will stay within the replay objective.
The replay path deserves its own test. A tailing consumer may read hot data from a broker cache, while a recovery consumer asks for older ranges that require object-storage fetches. The record’s schema identifier travels with the payload, but schema resolution may be served from a client cache, a registry cache, or a live registry request. Measure these cases separately:
- Hot read: record and schema are already cached.
- Cold read: the broker fetches an older range and the consumer resolves a schema it has not seen.
- Long pause: a consumer resumes after the registry client, broker cache, or credentials have changed.
- Cross-version reader: an old reader and a updated reader process the same retained range.
- Recovery read: a replacement broker and a restarted consumer rebuild state from shared storage.
The storage design affects when and where bytes are served, not the compatibility rule. KIP-405: Kafka Tiered Storage describes a local tier for active reads and a remote tier for older segments. KIP-1150: Diskless Topics describes a proposed Kafka direction in which object storage can replace broker disks as the primary durable location for topic data while local media can remain as cache or transient storage. Both proposals make the same operational point relevant to schema evolution: old records can be served through a path different from the one that accepted them.
Metadata is a separate concern. Kafka offsets and topic metadata, registry versions, and object-storage locations can have different consistency and retention rules. A replay runbook should identify the source of truth for each:
| State | Owner | Failure to test |
|---|---|---|
| Topic and partition offsets | Kafka metadata and consumer group state | Offset reset or group restoration points to a missing range |
| Schema definition | Schema Registry and its backing store | Version lookup fails or returns an unauthorized response |
| Object location | Shared-storage metadata and storage service | Range exists but index, manifest, or metadata lookup is delayed |
| Client cache | Serializer and registry client | Stale cache hides a policy or credential change |
| Rollout intent | CI/CD and deployment metadata | Consumer version and registry policy drift apart |
If the team cannot draw these ownership boundaries, an apparently harmless schema change can become a storage incident. The first alert may show deserialization errors, while the root cause is a registry endpoint, an object-store request path, or a missing old version.
3Failure, cost, and compatibility checks
A schema rollout is ready when its failure modes are observable and reversible. Use a worksheet with one row per reader and storage transition instead of a single “schema compatible” checkbox.
| Scenario | Measurements | Production gate |
|---|---|---|
| Updated producer, previous consumer | Decode errors, consumer lag, dead-letter records, and registry lookups | Old readers continue to process the agreed sample and alert thresholds |
| Updated consumer, retained old records | Cold fetch latency, schema lookup latency, and replay completion | The updated reader decodes a historical range from the actual serving path |
| Registry outage | Lookup failures, client cache hits, retry volume, and error age | The behavior for cached and uncached schemas is documented and tested |
| Registry version removal | Delete or compatibility-policy audit, lookup errors, and affected topics | Version deletion is blocked or approved only after retention and replay checks |
| Object-storage degradation | Fetch latency, throttling, retries, and consumer lag | Backpressure and retry behavior do not silently skip records |
| Broker replacement | Metadata rebuild, cache warm-up, cold-range reads, and decode errors | A replacement broker can serve retained data while consumers resolve schemas |
| Rollback | Producer version, consumer version, registry state, and offsets | The team can stop the rollout and resume the previous reader contract |
Cost follows the same paths. Keep registry operations, object-storage bytes, object-storage requests, network transfer, broker compute, cache capacity, and replay jobs as separate lines in the model. Long retention can increase the number of records that may need historical schema resolution, even when the registry itself stores only metadata. A replay that misses the broker cache may add object-storage reads and network transfer; use the provider’s current pricing and the workload’s actual read pattern rather than a remembered rate.
Compatibility testing should include the serializers and connectors that production uses. An Avro test does not prove Protobuf or JSON Schema behavior. A producer-only check does not cover Kafka Connect converters or a batch decoder. For each format and client family, keep an executable fixture with:
- the previous schema and its representative records;
- the candidate schema and the expected compatibility result;
- a updated reader and an old reader;
- a cold replay from the configured retention range;
- the registry authentication and failure behavior;
- the rollback command and the offset or replay boundary.
This is where a diskless design can improve the question being asked. Instead of assuming that a broker owns every retained byte, the team can test whether shared durable data, registry history, and reader state remain independently recoverable. The test still has to run. Shared storage changes the boundary; it does not supply the evidence.
4How AutoMQ changes the operating model
Once the neutral framework is clear, a Kafka-compatible shared-storage system can be evaluated against the same schema and replay contract. AutoMQ keeps Kafka protocol and ecosystem compatibility while replacing broker-local durable log storage with a Shared Storage architecture. The relevant question for schema evolution is where the encoded records, indexes, caches, and registry requests live during normal operation and recovery.
AutoMQ documentation describes compatibility as a deliberate storage-layer change around the Apache Kafka computing layer. Its Apache Kafka compatibility overview lists Kafka clients and ecosystem components as evaluation subjects. That makes Schema Registry integration a testable migration boundary: keep the producer and consumer fixtures, then run them through the target deployment and its actual registry endpoint.
The storage side is also explicit. AutoMQ’s architecture overview describes object storage as the primary repository and a WAL layer for the write path. Its WAL storage documentation is the place to verify the deployment’s selected backend and failure domain. Neither page replaces a schema test; they identify the storage paths that the cold-replay and broker-replacement tests must exercise.
The operating model changes in three useful ways:
- Durable bytes and brokers can be reviewed separately. A replacement broker may rebuild serving state from shared storage without treating its local disk as the system of record.
- Replay becomes a first-class workload. Historical reads, cache warm-up, object-storage requests, and schema lookups can be measured as one end-to-end path.
- Compatibility remains at the Kafka boundary. Existing producer, consumer, Connect, and Streams fixtures can be used to test schema behavior while the storage layer is evaluated independently.
These benefits depend on placement and configuration. Measure registry locality, private endpoints, identity and access management (IAM) permissions, cache capacity, object-storage retries, and network boundaries in the same environment that will carry production traffic. If the registry is outside the failure domain of the Kafka cluster, document how a zone or endpoint incident affects decoding. If schema history is replicated or backed up separately, test restoration together with offsets and retained records.
5Decision checklist and FAQ
Bring this checklist to a schema and storage design review:
- Who can read old data? Name current, paused, recovery, connector, and analytical consumers.
- Which compatibility direction applies? Tie backward, forward, full, or transitive checks to the rollout order and reader lifetime.
- Where is the schema source of truth? Record registry endpoint, authentication, backup, deletion policy, and client cache behavior.
- Where are the bytes? Identify broker cache, WAL or buffer, object-storage ranges, and the indexes needed for a cold replay.
- Which failures are gates? Test registry outage, version removal, object-storage degradation, broker replacement, and rollback.
- Which costs are visible? Separate registry calls, object-storage requests and bytes, network transfer, compute, cache, and replay jobs.
- What proves readiness? Keep a repeatable fixture with old and updated readers, retained records, a cold replay, and recorded stop conditions.
5.1Does diskless Kafka change schema compatibility rules?
No. Compatibility rules belong to the schema registry and the serializers that use it. Diskless storage changes how encoded records are retained and served, so the same rules must be tested against hot reads, cold replays, and broker recovery.
5.2Should the schema registry be stored in object storage with Kafka data?
That is a deployment decision, not a requirement of diskless Kafka. Keep the registry’s availability, backup, access control, and version-deletion policy explicit. A shared Kafka data path does not automatically make registry history available.
5.3Can a consumer decode a record after its broker cache is gone?
It can if the record remains readable from the durable path and the consumer can resolve the schema identifier. Test the complete cold-read path, including object-storage access, registry lookup, credentials, and client retry behavior.
5.4Does longer retention require more schema versions?
Retention does not create schema versions. It increases the period during which old versions may be needed by consumers, replay tools, and recovery jobs. Treat registry history and data retention as related contracts and set deletion gates accordingly.
5.5What should a pilot prove?
A pilot should prove the reader contract and the storage path together: registration and compatibility checks, old and updated consumers, long-pause replay, broker replacement, registry outage, object-storage degradation, authentication, observability, and rollback. Record the result for each serializer and consumer class that production uses.
A schema decision is ready when the team can trace one record from registration to replay without guessing which system owns the next step. The diskless question then becomes concrete: can shared storage serve the retained bytes while the registry and every reader preserve the same contract? Run the worksheet against a production-shaped workload, and start an AutoMQ evaluation with the compatibility fixtures, replay evidence, and rollback gates in hand.
