Table of Contents
Table of Contents
The most expensive part of a Kafka migration is often the first moment when the new cluster becomes responsible for a real business path. That is also the moment when teams discover that “compatible” did not cover a missing header, a different offset boundary, a consumer that commits too early, or an alert that watches the source but not the target.
A Kafka migration dry run moves that discovery earlier. For one week, keep the source authoritative, send a production-shaped copy of selected records to the target, run shadow consumers with no business side effects, and reconcile the observations from both sides.
“Costs almost nothing” describes the shape of the rehearsal, not a promise of free infrastructure. The plan avoids a second authoritative write path, reuses retained or mirrored records instead of inventing a synthetic workload, limits the test to representative topics, and gives the target only the compute and storage needed for observation. You still pay for the resources and transfer the test consumes. What you avoid is paying outage-sized risk to learn a cutover assumption was wrong.
1The rehearsal that belongs before the migration
Migration plans often jump from inventory to replication to cutover. The missing step is a bounded experiment that answers whether the target can behave like the source under the client and workload conditions that matter. A staging cluster can prove that a client connects, but not how a long-lived consumer group handles a rebalance against production-shaped records.
A useful rehearsal starts with a narrow statement:
The source remains authoritative for writes and business reads. The target receives a copy, serves isolated shadow reads, and returns evidence to a comparison process. No target result can acknowledge, commit, or trigger a critical side effect.
That boundary matters because a parallel run can quietly become a dual write. Once the target can acknowledge a producer or commit the production consumer group, the team has introduced a second authority. The test is now changing the system it was meant to observe.
Pick one migration wave with enough variety to expose assumptions:
- A topic with the normal producer path, including its message keys, headers, compression, and schema behavior.
- A consumer group that represents the dominant read pattern and its offset commit behavior.
- A less critical topic with a different retention or compaction policy.
- The production network, identity, ACL, and observability path.
Do not start by copying the whole cluster. The purpose is to make the target observable and the findings attributable.
2Shadow traffic architecture: mirror without committing
The safest arrangement has four separate roles: the source path, the replication or mirroring path, the shadow target, and the verifier. The source continues to serve normal producers and consumers. A replication bridge or controlled replay sends selected records to the target. A shadow consumer reads those records with its own group identity. A verifier compares stable evidence from both sides without becoming part of the production request path.
The target consumer must be isolated at more than the group ID. Its credentials should have only the permissions required for the test, its topic writes should be denied, and its output should terminate in a comparison sink or test store. If the application under test normally writes to a database, emits an alert, charges a customer, or publishes another event, the shadow path needs a stub, a dry-run mode, or a record-only adapter. “We will ignore the output” is not a control.
The verifier should compare facts that survive a cluster boundary. Depending on the workload, that may include a stable event ID, partition, timestamp, payload hash, schema identifier, and processing result. Kafka offsets are useful local positions, but an offset from one cluster is not automatically the same record boundary in another. Treat offset equality as a result to prove, not as the reconciliation key.
Keep acknowledgments on the source path. A producer may continue to receive the same source acknowledgment it received before the rehearsal, while the target catches up asynchronously. Likewise, a shadow consumer can read and measure without committing offsets used by the production group. This gives the team a real read path to observe without making the target a hidden participant in the business transaction.
3Five migration assumptions to verify in seven days
A week is long enough to see ordinary traffic variation, a deployment or restart, and at least one controlled fault drill. Track the same five assumptions every day, with an owner and an evidence link for each result.
3.1Client behavior survives the target boundary
Start with the contract closest to the application. Check bootstrap discovery, TLS or SASL settings, ACL evaluation, API versions, compression, message headers, schema handling, and the producer and consumer settings that affect retries, batching, fetches, and commits. Kafka protocol compatibility makes this a tractable test, but it does not make every deployment detail identical.
Record connected clients, exercised operations, observed errors, and settings that needed a target-specific change.
3.2The target receives the records the application thinks it receives
Compare record identity and shape, not only byte volume. Validate keys, headers, timestamps, partition placement, ordering within a partition, schema identifiers, tombstones, compaction behavior, and retention boundaries where those features matter. A target can show low replication lag while silently dropping a field that a downstream service uses.
Use a bounded sample and a repeatable reconciliation job that reports missing, duplicated, reordered, or transformed records with source and target positions. If the workload has no stable event ID, define a composite identity before the run.
3.3Shadow consumers behave under real read pressure
A shadow group should read the same topic class and encounter the same payload distribution as the production group. Observe fetch latency, consumer lag, rebalance events, deserialization errors, retry behavior, backpressure, and processing time. Keep its offset commits separate from production, and make the read window explicit so a restart does not replay an unbounded history.
Use the readings to find behavior that would make a cutover unsafe, such as a consumer that falls behind after a rebalance or an application timeout that appears only when the target serves a cold portion of the log.
3.4The operating controls work before the cutover
A migration is also an observability and ownership change. Confirm that dashboards distinguish source and target, alerts have target-side signals, quotas and throttles are visible, credentials can be rotated, and operators can tell whether lag comes from the replication path, the target, or the shadow consumer.
Exercise the actions the on-call team will need: pause and resume replication, stop a shadow consumer, rotate a test credential, add or remove a topic from the wave, and drain the target safely. If every action requires a specialist who will not be on the cutover call, the rehearsal has found an operating gap even if the records match.
3.5Failure and rollback boundaries remain legible
A dry run should answer what happens when the target is unavailable, the replication path falls behind, a shadow consumer restarts, or a target-side read produces a mismatch. Keep the source authoritative while you inject those failures. The expected result is a visible target-side defect with no change to production acknowledgments, production offsets, or critical external effects.
Write the rollback boundary in operational terms: stop mirroring, stop shadow reads, preserve the source, and retain enough evidence to explain the mismatch. If the real migration will later allow target writes, a separate cutover rehearsal must prove how target-only records and consumer progress are handled. A read-only dry run cannot prove that part, and saying so is a useful result.
4Fabricating realistic load without inventing a new workload
Synthetic load is useful for saturation testing, but it is the wrong default for a migration rehearsal. The risk comes from the shape of existing records and client behavior, so begin with records the system has already produced.
There are three low-risk ways to make the target see realistic work:
-
Mirror a selected live window. Replicate a small set of topics while the source remains authoritative. Keep the target retention bounded to the test window and apply a denylist for topics whose payloads or downstream effects cannot leave the source boundary.
-
Replay retained records. Read from a retained source window and write only to isolated target topics through a controlled bridge. Preserve keys, headers, timestamps, and partitioning where the tool supports them. Scrub sensitive payloads if the target account or test store has a different access boundary.
-
Re-run a captured client trace. Use a sanitized record and request trace to exercise a target client without connecting the application to production side effects. This is useful when live mirroring is too broad, but it must preserve the behaviors the cutover depends on.
In all three cases, measure the test's own footprint. Watch target storage growth, replication bandwidth, object storage requests, compute utilization, and log volume. “Almost nothing” is credible only when the team can show what it allowed the test to consume and why that amount was sufficient. It does not mean the cloud bill is zero.
Once the test needs a target that can carry the same Kafka-facing workload without pulling the production path into a second authority, architecture becomes part of the rehearsal design. AutoMQ is one Kafka-compatible target a team can evaluate in that role. Its Kafka compatibility documentation describes the client-facing boundary, while its Shared Storage architecture places durable stream data in object storage and reduces the coupling between broker compute and retained data.
That architecture can make the shadow target easier to size for a bounded experiment because the target does not have to become a second long-term home for broker-local data before the team can exercise Kafka clients and reads. It does not remove the five checks above. The target still needs client review, record reconciliation, consumer checks, operating controls, and failure evidence. Kafka compatibility is a reason to test the path. It is not the result of the test.
5The go/no-go gate
Do not end the week with a green dashboard and a meeting opinion. End it with a decision record that names the evidence, the unresolved assumptions, and the scope of the next wave.
A practical gate uses five rows, one for each assumption:
| Gate | Go evidence | No-go signal |
|---|---|---|
| Client behavior | Required clients connect and exercise their recorded operations | Unexplained errors, config drift, or unsupported behavior |
| Record fidelity | Reconciliation finds no unexplained missing, duplicate, or reordered records | Any mismatch without an owner and repair plan |
| Shadow reads | Consumer lag, fetch behavior, retries, and rebalances fit the agreed workload bounds | The shadow group falls behind or changes behavior under expected pressure |
| Operating controls | Dashboards, alerts, permissions, and pause or resume actions are exercised | An action or alert depends on an unowned manual step |
| Failure boundary | Target faults leave source acknowledgments, production offsets, and critical effects unchanged | The test can influence a production commit or side effect |
A go decision should be scoped. It can approve one topic class, one consumer group, or one migration wave. It should not silently approve every client, connector, schema, retention policy, and failure mode in the cluster. A no-go decision is also useful when it names the failed assumption, the owner, and the next experiment.
The follow-on cutover plan still needs authority boundaries, offset mapping, consumer handoff, and rollback for target-side writes. The blue-green Kafka migration guide covers those acknowledged-message and authority questions. For a broader dual-cluster evaluation, use the dual-cluster validation guide. The dry run comes first: it tells you whether the target deserves a cutover plan.
A migration rehearsal earns its value by leaving the production path boring. At the end of the week, the source should remain authoritative, the target should have a bounded evidence trail, and the team should know which assumptions are proved. With selected records, isolated consumers, and controlled capacity, almost nothing has to change before a larger wave.
To run this rehearsal against a Kafka-compatible target, start an AutoMQ evaluation with one representative topic class, its client matrix, and the gate above. Keep the experiment bounded, then decide from the evidence.
6References
- Apache Kafka documentation, including producer, consumer, replication, and topic configuration references.
- Apache Kafka replication design, for leader, follower, and in-sync replica behavior.
- Apache Kafka consumer configuration, for group identity, offset commits, and fetch behavior.
- AutoMQ compatibility with Apache Kafka.
- AutoMQ architecture overview.
- AutoMQ migration overview.
7FAQ
7.1What is a Kafka migration dry run?
It is a bounded rehearsal in which selected production-shaped records reach the proposed target, isolated consumers read them, and a verifier compares the results while the source remains authoritative.
7.2What is Kafka shadow traffic?
Kafka shadow traffic is a copy of selected records or a replayed record window that a target-side consumer reads for observation. The shadow path uses separate identities and is prevented from acknowledging production writes, committing production offsets, or triggering business side effects.
7.3Can Kafka clusters run in parallel during a migration?
Yes, but parallel operation needs explicit ownership. Keep one source of truth for production writes and production consumer progress during the rehearsal. Give the target its own topics or replication boundary, consumer groups, permissions, dashboards, and stop conditions.
7.4Does Kafka protocol compatibility make a migration safe?
No. It reduces the application-surface risk that the rehearsal needs to investigate, but teams still need to validate security settings, record fidelity, offsets, consumer behavior, quotas, observability, and rollback boundaries against their workload.
7.5Does “costs almost nothing” mean the dry run is free?
No. It means the design limits incremental exposure by using selected traffic, bounded retention, isolated reads, and no second authoritative business path. Measure the target resources, storage, bandwidth, and operations the experiment consumes before approving a larger wave.
