Table of Contents
Table of Contents
Raw JSON is a productive place to start an Apache Kafka® stream. A producer can add a field without waiting for code generation, a developer can inspect a record with a terminal, and a consumer can ignore fields it does not understand. The trouble arrives later, when customer_id appears as both a number and a string, created_at carries two time conventions, and a missing field means three different things to three different consumers.
That is when a payload format becomes an operating problem. The topic still accepts bytes, but the teams around it have lost a shared answer to a more useful question: what does this record promise, and which versions can coexist during a rollout? Moving to Avro or Protobuf can restore that answer, provided the migration treats serialization as a client contract with a controlled transition. The work is a staged change to producers, consumers, schema tooling, and topics. The Kafka broker can keep doing its transport job throughout.
1JSON is flexible until the fields start lying
JSON's early advantage is real. It keeps the producer and consumer boundary loose while a team discovers the event shape. That flexibility becomes expensive when the event is reused outside the service that first published it. A field name starts carrying undocumented assumptions about units, nullability, identity, and time. Consumers add defensive parsing, default values, and special cases, and those local fixes become the actual contract because nobody can point to a versioned definition.
The failure is rarely a parser exception on the first day. A consumer may parse the record successfully and still make a wrong decision because "amount": 25 means dollars in one producer and cents in another. A replay job may read a retained record after the original service has changed its interpretation. A sink may accept the field but map a string timestamp differently from a numeric timestamp. The stream stays available while the meaning drifts.
A schema gives the team a named, reviewable boundary. Avro carries a writer schema that can be resolved against a reader schema, while Protobuf assigns stable field numbers and defines rules for adding or reserving fields. A schema registry can store versions and apply compatibility checks, but it does not decide whether amount still means the same business quantity. That semantic decision belongs in the event contract and its review process.
Format choice should follow the record's life. Avro is often a practical fit for event contracts that need schema resolution and compact binary records. Protobuf is a strong fit when generated types and stable field numbers are already part of the development model. JSON Schema can preserve a text-oriented workflow when human-readable payloads matter. A useful first inventory has four columns: the fields that exist, the meanings consumers rely on, the writer and reader code that handles them, and the retained history that must remain readable. Mark every field as required, optional, deprecated, or semantic-risk. Identify the consumers that parse the payload directly, the consumers that use a shared library, and the jobs that replay old offsets. This inventory is the baseline for a migration that can be tested instead of inferred.
Illustrative scenario: An order event has lived as raw JSON long enough for three services to depend on it. The producer team wants a typed contract, while one analytics consumer still replays retained events and a notification service is deployed on a slower release cycle. The migration succeeds when those facts shape the rollout sequence. It does not require an invented outage, customer name, or performance claim.
2Staged by design: parallel first, flip later
The safest default is to keep the legacy topic readable while a parallel topic carries the versioned record. A producer publishes the same business event to both paths for a controlled period, or a translation service reads the legacy topic and writes the versioned topic when producer changes must remain isolated. The exact fan-out mechanism depends on delivery and ordering requirements, but the boundary should be explicit: one topic has a legacy wire format, the other has a declared schema and subject.
A parallel topic costs some operating attention, yet it removes a dangerous ambiguity. A mixed topic can contain JSON, Avro, and Protobuf bytes that all look like valid Kafka records to the broker but require different deserializers at the consumer. A topic-level contract lets a consumer choose one decoding path and lets the platform observe progress by topic. It also gives the team a clean place to stop publishing after the cutover without rewriting the historical contract in place.
Use a sequence with a visible gate at each transition:
- Describe the legacy contract. Capture field meanings, key behavior, null rules, timestamps, retention, partitions, and known consumers. A schema file without these details is an incomplete migration input.
- Register the candidate schema. Choose a subject naming rule, compatibility mode, owner, and failure behavior. Test representative old records against the candidate reader and candidate records against the intended readers.
- Create the parallel topic. Copy the partitioning key, ordering assumptions, retention policy, access controls, and operational alerts. Record whether the topic is populated by dual-write code or a translation path.
- Run consumers in shadow or dual-read mode. Compare decoded fields, keys, timestamps, business decisions, and error handling. A consumer that can deserialize a record has passed one gate, not the whole migration.
- Flip one consumer at a time. Keep rollback tied to a measurable signal, such as decode failures, business validation failures, lag, duplicate side effects, or replay mismatch. Stop the flip when the signal crosses the agreed threshold.
- Retire the legacy write path after the read window. Keep the old topic and its decoder until every replay, backfill, and rollback obligation has expired. Delete it only after the owner can prove that no supported workflow still needs it.
3Writing schemas without stalling producers
A schema migration stalls when registration becomes a release ceremony with no clear contract. It moves at a controlled pace when the candidate schema is tested against actual writer and reader behavior before it reaches production traffic. The producer team should own the record it publishes, while the platform team supplies the registry, client libraries, CI checks, and operational signals that make the contract enforceable.
Start with the fields that cause the most ambiguity. Give each one a type, unit, null rule, default behavior, and semantic owner. Preserve a stable event key. Keep event time distinct from ingestion time. If a field cannot be defined precisely, mark it as a migration decision rather than disguising the uncertainty with a broad string type.
For an additive change, an optional field with a documented default can help older and current readers coexist. A rename needs more care because it may be a structural change, a semantic change, or both. Preserve an alias where the format supports it, or write both names during a deprecation window. Removing a field requires evidence that retained readers, replay jobs, and downstream exports no longer depend on it. Changing the meaning under the same field name is a contract change even if the serializer accepts it.
The Apache Avro specification and the Protobuf field-update guidance describe format-specific rules that belong in the compatibility test suite. Those rules are necessary, but production acceptance also needs fixtures from the legacy topic. Test the candidate consumer against old records, the legacy consumer against proposed records when the rollout requires it, and both consumers against malformed or incomplete input. A registry check should be one gate in that matrix, not its substitute.
The registry itself is part of the migration surface. Record its endpoint, authentication, subject naming convention, compatibility mode, schema ID handling, client library version, and outage behavior. Decide whether producers fail closed when registration is unavailable or use a cached schema under a documented policy. Decide how consumers behave when they encounter an unknown ID or a record from outside the supported window. These are operating decisions, not details to discover during cutover.
A format change can also change record size, CPU work, and error visibility. Measure those effects with the workload that matters to the service, then keep the result as a migration observation rather than a universal promise. Binary encoding may reduce repeated field names, while generated code may move validation earlier. The broker still transports the serialized bytes. Producers, consumers, and their SerDes determine the format work.
4Dual reads and the legacy-data window
A dual-read consumer reads the versioned topic through the versioned deserializer and continues to read the legacy topic through the old path until the window closes. Some systems use a single consumer application with two input paths; others run a shadow consumer group that compares normalized records before the production group moves. The design should preserve keys, event identity, ordering expectations, and side-effect rules so that comparison does not create duplicate business actions.
The old-data window has two separate jobs. It protects rollback, because a consumer can return to the legacy path while the old topic remains authoritative for a defined period. It protects replay, because retained JSON records may still be needed for backfills, audits, or rebuilding a downstream view. Confusing these jobs leads to premature deletion. A consumer may be fully cut over while a replay workflow still depends on the legacy decoder.
Define the window from obligations rather than a convenient calendar date. Ask when the slowest supported consumer will be upgraded, how far back a supported replay may start, how long a rollback remains credible, and whether an external export still preserves legacy bytes. The answer may be different for different topics. A high-value event with long retention needs a different retirement test from a short-lived operational signal.
A migration scoreboard should compare the two paths at the record level where possible. Useful checks include:
- Identity: the Kafka key, event identifier, and partitioning decision remain aligned.
- Shape: required fields, types, defaults, aliases, and unknown-field handling match the contract.
- Meaning: timestamps, units, enum values, and null behavior produce the same business interpretation.
- Progress: consumer offsets, lag, retry counts, dead-letter records, and replay position stay within the planned range.
- Side effects: the versioned path does not send a notification, charge an account, or write a state change twice during comparison.
If the paths diverge, keep the record and schema version that exposed the difference. A count of successful deserializations will miss a valid but misinterpreted event. The most valuable evidence is a small, reviewable fixture that a producer, consumer, and contract owner can inspect together.
5What changes when the broker stays out of the format choice
The serialization migration belongs at the client and compatibility boundary. The Apache Kafka protocol documentation covers the broker-facing request surface, while serializers, deserializers, schema registries, and application code define how payload bytes become structured data. Kafka transports records and coordinates topics, partitions, offsets, and consumer groups. That separation means a team can change JSON to Avro or Protobuf without requiring the broker to understand the schema, as long as the client path and topic contract are controlled.
This is also the boundary at which AutoMQ fits the migration. AutoMQ is Kafka-compatible, so the relevant migration work stays in producers, consumers, SerDes, registry integration, and compatibility tests. Its Kafka compatibility documentation defines the client-facing boundary, while the Shared Storage architecture overview describes a storage model underneath the Kafka data plane. Neither changes the ownership of a schema contract.
That distinction matters during a platform migration. A Kafka-compatible broker can reduce broker-side rewrite work, but it does not copy schema subjects, serializer settings, registry credentials, application defaults, or a legacy read window by itself. Test those assets as client behavior. The format migration and the storage-platform migration may share a cutover plan, yet they remain separate gates with separate rollback actions.
For teams evaluating AutoMQ, the practical question is whether the existing producers and consumers can keep their Kafka client and SerDe assumptions while the data plane changes underneath. The answer should come from a workload-shaped compatibility test that includes the two topic paths, registry behavior, retained records, consumer groups, and operational tooling. The production guardrails for Schema Registry runbooks and schema test environments for event-driven development cover adjacent controls for making that evidence repeatable.
6A production migration checklist
The migration is ready for production when the team can answer each question without relying on the person who first wrote the JSON producer. Use the checklist as a release gate and attach the evidence to the topic, schema subject, and application change record.
6.1Before the first dual write
- Scope: The legacy topic, parallel topic, key, partitions, retention, access policy, and owning team are recorded.
- Contract: Every migrated field has a type, semantic meaning, unit, null rule, default, and retirement status.
- Registry: Subject naming, compatibility mode, authentication, schema ID behavior, and outage policy are tested.
- Fixtures: Representative legacy records, proposed records, malformed records, and replay records are stored in version control or an approved test store.
- Consumers: Direct parsers, shared SerDes, Connect converters, stream jobs, exports, and replay tools are inventoried.
6.2Before each consumer flip
- Dual read: The consumer can read both formats without mixing deserializers on one topic path.
- Comparison: Keys, identity, timestamps, business meaning, lag, errors, and side effects are compared.
- Rollback: The previous consumer path, deployment artifact, offset plan, and decision owner are ready.
- Observability: Decode failures, schema lookup failures, validation failures, dead-letter records, and replay mismatches have alerts or review signals.
- Load: The test covers the production-shaped payload size, fan-out, retention access, and failure handling that matter to the service.
6.3Before retiring JSON
- Writes stopped: The legacy producer or translation path has a clear stop condition and no untracked writers.
- Reads covered: Supported consumers, backfills, audits, exports, and rollback paths have either moved or reached their retirement condition.
- Window closed: The owner has recorded the last supported legacy offset or event time and the reason it is safe to remove the old decoder.
- Recovery tested: A current consumer can rebuild its state from the versioned topic, and the team knows how to preserve the legacy copy if the test fails.
- Documentation updated: Topic ownership, schema version, client configuration, runbook, and deprecation decision point to the same record.
The migration nobody regrets is rarely the one with the most elaborate schema. It is the one that leaves no ambiguous middle state. Producers know which topic they write, consumers know which versions they read, the registry enforces the intended compatibility direction, and the old bytes remain available until their real obligations end.
If your team is planning a Kafka-compatible platform change at the same time, run this format migration as its own client-side gate. Start an AutoMQ evaluation with the producer, consumer, registry, parallel topic, and retained-data window in the test plan. The useful result is not a claim that every migration looks the same. It is evidence that your event contract survives the transition.
