Table of Contents
Table of Contents
A Kafka migration can look healthy while it is still unsafe to cut over. Producers may be writing to the destination, consumers may be reading a few test topics, and the dashboard may show green checks. None of that proves that offsets, replay behavior, storage reads, network paths, and rollback decisions will hold when production traffic moves.
A diskless Kafka migration makes that gap easier to miss because the durable data path changes at the same time as the cluster boundary. The broker still handles Kafka requests and partition leadership, but retained stream data lives in shared object storage, with caches and write buffers between clients and that durable layer. A production rehearsal therefore has to test contracts between clients, metadata, storage, network, and operators. The useful output is a set of measured gates, not a confident feeling about a successful demo.
1What a migration rehearsal has to prove
A rehearsal is a production-shaped exercise that can be stopped without harming the source cluster and repeated with the same evidence. It is different from a compatibility smoke test, which usually proves that a client can connect and produce a record. It is also different from a data copy, which proves that bytes arrived somewhere but may say nothing about consumer positions or the path used to serve a replay.
Use four contracts to define the rehearsal before choosing a migration tool:
- Data contract: acknowledged records arrive at the destination with the expected ordering, retention, compaction, and transactional behavior for the topics in scope.
- Client contract: producers, consumers, Kafka Streams applications, Kafka Connect tasks, authentication, quotas, and monitoring integrations behave as they do against the source cluster.
- Operations contract: the team can observe replication, lag, cache behavior, object-storage requests, network errors, and broker recovery using named owners and clear thresholds.
- Rollback contract: a failed cutover has a known stop condition, an offset strategy, and a safe way to resume service on the source without losing acknowledged writes.
The rehearsal should end with a decision for each contract: pass, fail with a documented action, or out of scope for this migration wave. “We can migrate” is too broad to be an operational result.
2Start with a baseline you can replay
A destination cluster cannot be judged against a source cluster if the source behavior was never recorded. Capture a baseline during a representative window and keep the capture tied to the workload, not to a generic cluster average. The baseline should include producer rate and record size, partition-level append rate, consumer positions, processing time, rebalance duration, request latency, and the storage and network signals that explain those measurements.
Record the boundaries that can change during migration. These include client library versions, security protocol and identity provider, topic configuration, partition count, replication or durability settings, quotas, compression, message format, and the cloud endpoint used for object storage. A rehearsal that changes several of these at once cannot tell you which change caused a result.
A useful baseline has two workload shapes. A tailing workload reads data close to the producer’s current position and exercises the hot path. A replay workload starts from older offsets and exercises cache misses, object-storage reads, fetch sizing, and consumer processing under catch-up pressure. Both are needed because a migration can pass tailing tests and fail when a recovery job replays a long range.
Keep a simple evidence table for every run:
| Evidence | Source baseline | Destination rehearsal | Gate question |
|---|---|---|---|
| Producer acknowledgements and error rate | Capture by topic and partition | Repeat with the same client settings | Are acknowledged writes accepted consistently? |
| Consumer position and freshness | Capture position, processing time, and commits | Compare with the same consumer group behavior | Can consumers make progress without an offset gap? |
| Fetch latency and response bytes | Record p50/p95 or the selected service SLO | Measure tailing and replay separately | Did the read path change under the same workload? |
| Storage and network signals | Record request latency, errors, bytes, and endpoint path | Attribute the same signals to the candidate design | Which dependency owns a delay or retry? |
| Recovery actions | Time and steps for restart or rebalance | Repeat the same drill | Can on-call staff follow the runbook without improvising? |
The table is intentionally about evidence, not a target number. Set thresholds from the freshness and recovery objectives that the application owners already accept. If no owner can state those objectives, the migration is not ready for a cutover gate.
3Dual-run means more than copying records
A dual-run phase keeps the source serving production while the destination receives a controlled copy or replication stream. The destination should be exercised by the same client behaviors that matter after cutover, even when only a subset of topics is in scope. Create a topic and consumer-group inventory first, then classify each item by migration order and rollback sensitivity.
During dual-run, measure three distances instead of one replication-lag number:
- Write distance: the difference between the source’s acknowledged end and the destination’s replicated end for each partition.
- Read distance: the difference between the destination data available to a test consumer and the position that consumer has validated.
- Decision distance: the work still required before the team can stop writing to the source or reverse the change.
The third distance is where many migrations lose control. A destination may be caught up while a connector, schema registry client, or operational alert still points at the source. Treat those dependencies as part of the dual-run contract. A migration is not ready because the bytes are present; it is ready when the services that interpret those bytes can use the destination under the intended identity and policy.
Run a controlled replay during dual-run. Pick a bounded offset range and validate record count, key ordering where the application depends on it, headers, timestamps, tombstones, and transaction markers where applicable. Keep the source and destination evidence side by side. If the destination uses a different storage path, record whether the replay was served from a cache, a write buffer, or object storage. This distinction becomes important when a recovery run starts cold.
4Design the cutover as a sequence of gates
A cutover should be a short sequence of reversible actions with an explicit owner for every gate. Avoid a single “switch traffic” command that hides several decisions. The following order keeps client behavior, offsets, and data validation visible:
- Freeze the scope. Stop topic creation, configuration changes, and unrelated deploys for the migration wave. Record the source offsets and the exact client versions in the run log.
- Confirm destination readiness. Check replication distance, topic settings, authentication, quotas, observability, storage permissions, and network routes. A passing data check cannot compensate for a missing alert or an untested identity.
- Move consumers with a known position. Redirect one low-risk consumer group first, validate its position and processing output, and keep the source read path available for rollback. Record the time between the last source commit and the first destination commit.
- Move producers in controlled batches. Roll clients in a defined order. Watch acknowledgements, retries, duplicate handling, and transaction outcomes while the source remains available for the rollback policy.
- Validate before decommissioning. Compare partition-level end offsets, consumer positions, sampled records, and application-side results. Keep the source cluster intact until the rollback window expires and owners sign the gate.
The order matters because producers and consumers do not fail in the same way. A producer can receive an acknowledgement while a consumer still lacks the data path or credentials needed to read it. A staged cutover makes that mismatch visible while the source remains a recovery option.
5Exercise rollback before you need it
Rollback is a tested state transition, not a sentence in a change plan. Choose a failure point before the rehearsal starts, then trigger it with a non-critical migration wave. Useful drills include object-storage request errors, a broker restart during a replay, a consumer rebalance during producer rollout, a network route failure, and a validation mismatch on a sampled partition.
For each drill, answer four questions in the runbook:
- What signal stops the rollout?
- Which writes are acknowledged at that point, and where are they durable?
- How will consumers resume from a known position without skipping or duplicating records beyond the application’s tolerance?
- Which operator owns the decision to resume, retry, or return to the source?
A rollback path that depends on reconstructing offsets from a dashboard is not a rollback path. Keep an immutable run log with source and destination offsets, client rollout order, and the exact time of each gate. If the migration mechanism provides write forwarding or another method to preserve writes during a reversal, test its failure behavior and limits; do not infer safety from the feature name.
6Check the storage model behind the word “diskless”
“Diskless Kafka” does not describe one implementation. Traditional Apache Kafka® keeps an active log on broker-local storage. Kafka Tiered Storage adds remote storage for older segments while retaining a local active tier; KIP-405 documents that model. The diskless topic proposal in KIP-1150 is a useful reference for the direction of the Apache Kafka discussion, but a proposal is not proof that the release and implementation in your plan support every client and workload you run.
A shared-storage design moves the durable ownership boundary. Brokers still handle the Kafka protocol, partition leadership, scheduling, and group coordination, while stream data is retained in shared object storage. Memory caches and a write-ahead log can serve hot data and acknowledge writes before background upload, depending on the implementation and WAL type. That changes what a migration rehearsal must observe:
- Cache: Which reads are hot, which are catch-up reads, and how eviction affects replay latency?
- Metadata: Can a replacement broker discover partition and stream state without copying a local log directory?
- Object storage: Are endpoint placement, request limits, retries, and permissions visible in the same dashboard as consumer progress?
- Network: Do client, broker, and object-storage paths cross zones or private-network boundaries that change cost or failure behavior?
The rehearsal should compare these paths under the same workload. Do not compare a warm-cache tailing test on one architecture with a cold replay on another and call the difference a storage result. The Kafka compute-storage separation and Tiered Storage comparison is useful background when the team needs to make this distinction explicit.
7How AutoMQ fits after the neutral gates
Once the rehearsal identifies storage ownership, recovery behavior, and client compatibility as the decision points, a Kafka-compatible shared-storage system becomes a concrete candidate. AutoMQ keeps the Kafka protocol boundary while using S3Stream and S3-compatible object storage for durable stream data. Its architecture documentation describes the broker, metadata, cache, WAL, and object-storage path that the same rehearsal measurements can exercise.
This matters for migration because compute and retained data have different lifecycles. A broker replacement or capacity change can be evaluated as a compute operation when partition data is not tied to that broker’s local disk. The team still has to validate metadata recovery, cache warm-up, WAL behavior, object-storage permissions, and network placement. Shared storage removes one class of data movement; it does not remove the need for a recovery drill.
AutoMQ also documents a Kafka Linking workflow for migration and replication. Treat it as a mechanism to test against the four contracts above, not as a reason to skip them. Validate source-to-destination replication, staged client movement, offset continuity, application output, and the rollback procedure with the client versions and topic features in your scope. The AutoMQ migration and replication page describes the product workflow; your rehearsal remains the evidence that it fits your environment.
8Production decision checklist
Before approving a migration wave, ask the owners to sign the checks that matter to their service:
- Compatibility: Have all producer, consumer, Streams, Connect, security, quota, transaction, compaction, and monitoring paths been exercised on the destination?
- Data validation: Have sampled records, offsets, ordering rules, tombstones, and application-side results been checked during dual-run and after cutover?
- Recovery: Has the team repeated broker, cache, object-storage, network, and consumer-rebalance drills with an explicit stop and resume action?
- Cost model: Are object-storage capacity, request, compute, network, and observability lines separated using the selected provider’s current pricing and endpoint topology?
- Rollback: Can the team state the last safe source offset, the handling of acknowledged writes, and the owner who can stop the rollout?
- Scope: Is the next wave small enough to fail without turning a rehearsal into an incident?
8.1Is a diskless Kafka migration only a replication job?
Replication moves data, but migration changes where clients write, where consumers commit, how operators observe failures, and how the team returns to the source. The rehearsal must cover those state transitions and the storage path that serves hot and replay reads.
8.2Does Tiered Storage make a rehearsal unnecessary?
No. Tiered Storage and diskless Kafka can both use object storage, but they place durable data differently. A local active tier, a remote segment tier, and a shared durable stream create different cache, recovery, and network tests. Confirm the implementation and release you plan to operate.
8.3What should be the first migration wave?
Choose a topic and consumer group with a known freshness objective, representative record shape, and a rollback owner. Keep the wave small enough to repeat, then expand only after the same evidence gates pass under tailing, replay, and failure workloads.
A migration rehearsal earns its name when the team can explain what happens to a record, an offset, and a recovery decision at every gate. Start with the baseline, force the failure, and keep the rollback path concrete. If a shared-storage Kafka design fits those measurements, start an AutoMQ evaluation using the same workload and stop conditions.
