Table of Contents
Table of Contents
The migration dashboard shows every topic in green. Records are landing on the destination cluster, replication lag sits inside tolerance, and then the incident queue lights up anyway. A payments consumer group resumes from an offset that no longer points at the same record. A sink connector keeps writing to the warehouse from the source cluster because its connection string was never re-pointed. An enrichment service starts decoding messages with a schema ID that the destination registry has never seen. The topics moved. The applications did not.
Most migration planning is built around the part that already works. Topic mirroring, partition mapping, and byte-level copying are solved problems with mature tooling, so they become the plan's spine. The part that fails is the wiring around each topic: the consumer groups that hold offsets in the source cluster's coordinate system, the schema registry that turns bytes into agreed record formats, and the connectors that carry side effects outside the cluster. Moving Apache Kafka without moving those attachments is like moving a data center but leaving every cable in the old building, and the attachments snap back the first time they are asked to do real work. That pull is data gravity, and the migration math follows. Bytes can be replayed while a miswired consumer cannot, so a handful of attachments cost more than the history they read. Treat downstream systems as the scarce resource, plan the cutover around them, and the topic movement becomes the quiet part.
1Moving topics is not the migration
Copying a topic is a logistics problem. Moving the readers of that topic is a coordination problem, and the two obey different math. A topic can be mirrored in the background for as long as retention allows, and a replication job often finishes early with nothing left to do. The attachments do not finish early. Each one has a position, a format, or a side effect that exists somewhere else, and each one has to be re-anchored at a specific moment or it fails in production.
Migration effort is not predicted by topic count, or even retained volume. It is predicted by the number of things that remember which cluster they belong to. A topic does not remember; a consumer group does, a schema registry does, and a connector does, and they remember in different places:
- Consumer groups store committed offsets in the source cluster's internal metadata. When a group moves to a destination cluster, it resumes from whatever offset the destination knows, which may be a different record, a missing record, or the earliest offset available.
- Schema registries hold a global ID per schema, and producers embed that ID in each message they serialize. A consumer that reads with the destination registry must get the same ID for the same schema, or deserialization fails at the edge, after the brokers have done their job.
- Connectors keep their own offsets, plugins, and credentials, and they produce side effects that a broker never sees: rows in a warehouse, objects in cloud storage, records in a downstream database. Re-pointing a connector is not copying its data, and getting it wrong can double-write or silently stop the sink.
A migration plan that counts topics measures transport capacity, not migration. The real plan runs in consumer groups, schema versions, and connector instances, and those are the rows that belong at the top.
2What data gravity actually consists of
Data gravity is a cloud-era metaphor, but in Kafka it has a precise, mechanical meaning. Gravity is whatever makes a topic stay where its consumers already are. Every consumer group committing to the source, every producer serializing against a registry, every connector draining or feeding the topic adds another object to re-root before the move is clean. The topic grows heavy with commitments, not bytes, and the commitments take three shapes.
| Attachment | Where its memory lives | How it fails if forgotten |
|---|---|---|
| Consumer groups | Committed offsets in source cluster metadata | Resume at the wrong offset, replay, or stall |
| Schema registry | Global schema IDs written into message bytes | Deserialization failures for every encoded message |
| Connectors | Connector offsets, plugins, and credentials | Duplicate writes, stopped sinks, external side effects |
The table hides the uncomfortable part: timing. A consumer group fails at cutover, when it tries to resume. A schema mismatch fails at decode time, often in a consumer that was migrated weeks earlier and had been reading fine from cached state. A connector can run for a month without anyone noticing, because its sink still accepts writes and only the source of truth has gone stale. The same attachment fails at a different speed in a different direction, so no single cutover checklist catches all three. There is only a planning order that surfaces each attachment before the failure does.
Gravity also hands you a diagnostic. When a migration feels stuck, the topic copy is almost never what is stuck. The questions pile up downstream: which group owns the partition map, who approves the serialization change, where the connector's secret lives now. Follow them and the migration becomes legible; ignore them and every meeting replays the same three questions in a new order.
3Reverse-order planning: downstream first
Conventional migration planning starts at the source and works forward: mirror topics, watch the lag, flip producers, move consumers, fix whatever breaks. The cutover point sits at the end of the path, understood as the moment the bytes arrive. The downstream-first plan starts at the opposite end. Fix the cutover point first, at the readers and the sinks, then push the data movement back until it meets them.
Order matters because the cutover point is where every attachment converges at once. If attachments are not re-anchored before the data arrives, cutover day becomes a scramble to re-point consumers, registries, and connectors inside one window. Keep the cutover separate from the data movement and the scramble disappears. The sequence reads in reverse for a reason:
- Inventory every consumer group, schema dependency, and connector on the topics in scope, the same way you would inventory an estate, not a cluster.
- For each attachment, define its future coordinate before any records move: the destination offset map for every group, the destination registry and matching schema IDs, the destination connector endpoints with a duplicate-or-gap decision documented for each one.
- Start the dual write only after those coordinates exist, so the destination fills in the background while downstream systems keep reading the source as normal.
- Flip the attachments first, in dependency order, while the data has had time to catch up, then move producers last, once nothing downstream is still looking for the source.
The riskiest work stops being a cutover event and becomes a planning exercise. Cutover is short because the attachments were always the plan. Data movement is long and safe; re-wiring is early, and that is where failures live.
Downstream-first planning also sets a measurable requirement for the destination. If each consumer must change its bootstrap string, group ID, or deserializer to reach the destination, then each group is a change ticket, and the migration recreated the gravity it meant to escape. The destination must let consumer groups keep their protocol behavior and offset semantics while the data plane changes underneath. Check that property before estimating how many attachments must be re-wired rather than re-pointed.
This is the entry condition that puts AutoMQ in scope. AutoMQ is a Kafka-compatible streaming platform that keeps the Kafka protocol, so consumer groups, clients, and tooling keep their protocol behavior and offset semantics while the storage underneath changes. Its compatibility with Apache Kafka and architecture overview are the two references to read while scoring targets. The point is not that a compatible destination makes gravity disappear; it is that gravity concentrates where the real decisions are: schema and connector planning instead of a rewrite of the reader fleet.
4Re-wiring connectors and schemas
With consumer groups held steady by protocol compatibility, what remains is the two attachments that carry external state of their own, and they re-wire on different schedules set by what each embeds on the wire. Schemas move first, because they are part of the message. A producer serializes against a registry and writes the schema ID into each payload, so that registry must exist at the destination before the first encoded record arrives. Register the schema set in the destination registry, confirm IDs resolve to the same schemas on both sides, and only then let producers targeting the destination write. If the registries disagree, moving topics first lands the disagreement inside consumers, where it is most expensive to find and fix.
Connectors move last, because each connector is a small migration of its own, with its own offsets and its own failure direction:
- A source connector feeding the topic can be duplicated against the destination, with a documented decision about which copy wins during overlap.
- A sink connector reading the topic re-points only after the destination has caught up enough that the sink sees no gap, then disables on the source once the flip is verified.
- A connector's configuration, secrets, and internal offsets travel as a unit: they pin its position as firmly as a consumer group's committed offsets do.
The order is schemas before producers, connectors after the data has caught up, and consumer groups held steady throughout by protocol compatibility. It is an ordering discipline, not a tool feature, and it can be written down regardless of which replication tool moves the bytes.
5A migration boundary table
Every migration needs a visible statement of what moves with which side. Without one, cutover day is an argument about ownership: does the warehouse sink belong to the cluster project, or to the data platform? The boundary table settles the argument ahead of time by assigning each concern to a side and writing the rule that keeps it consistent.
The boundaries that usually emerge: topic bytes belong to the destination once dual write proves parity; consumer offsets belong to the groups and resolve through a mapping; schema IDs belong to whichever registry the producers of record use; connector side effects belong to the connector owner, rollback path included. None of these lines are exotic. What makes the table work is that the lines exist before the data moves, so nobody renegotiates ownership from inside a cutover window.
That closes the loop. The dashboard that opened this plan shows green topics and still ends in incidents, because it measures the wrong thing. It measures where the bytes are. A migration is done when the attachments answer the question that actually matters: which cluster do you belong to now? When every consumer group, schema, and connector has a written answer and a rollback position, the topic movement finishes quietly behind them. Work through the same downstream-first planning against the AutoMQ Open Source project, and start with the inventory of attachments, not with the topics. The list of things that remember is the migration.
6References
- Apache Kafka documentation
- Apache Kafka consumer design, including consumer groups and offsets
- Apache Kafka Connect
- Apache Kafka replication
- AutoMQ compatibility with Apache Kafka
- AutoMQ shared streaming storage overview
- AutoMQ migration overview for Kafka-compatible cutovers
- Kafka migration's hardest part is moving the clients
- Beyond MirrorMaker 2: Kafka migration with zero downtime
- Dual-cluster validation for low-risk streaming migration
7FAQ
7.1What is data gravity in a Kafka migration?
Data gravity is the pull that keeps a topic attached to the systems that already read it. In a Kafka migration it names the real cost: consumer groups with committed offsets, schema registries embedded in message formats, and connectors with their own offsets and side effects. Topics copy cleanly in the background; the attachments have to be re-rooted at a specific moment.
7.2Why do consumer groups matter more than topics?
A topic is a byte stream that can be mirrored for as long as retention allows. A consumer group remembers its last committed offset in the source cluster's coordinate system, and on the destination that position may not correspond to the same record. Consumer groups turn a mirroring job into a migration: they are the first attachment to fail at cutover.
7.3What is the correct order for moving schemas and connectors?
Schemas move first, because the registry ID is embedded in each encoded record. The destination registry must hold the same schema set and resolve the same IDs before producers target it. Connectors move last, with a documented duplicate-or-gap decision per source and sink, and each travels with its own offsets, plugins, and credentials.
7.4How do I map consumer group offsets from source to destination?
Build a partition mapping between the clusters, snapshot each group's committed offsets on the source, and translate them through the mapping into destination offsets. Run that check on a schedule against a written lag tolerance so a parity break is caught before the reader fleet flips.
7.5Does choosing a Kafka-compatible destination change the plan?
It reduces one part. A Kafka-compatible destination that preserves the protocol lets consumer groups keep their offset semantics and bootstrap behavior while storage changes underneath, so the reader fleet is re-pointed rather than rewritten. The inventory still applies, and the remaining work concentrates on schema and connector planning.
