Blog

When Not to Migrate: Kafka Projects That Should Stay Put for Now

Table of Contents

Table of Contents

It starts with a reasonable question. A renewal quote lands high, an incident drags past its budget, or an executive reads a vendor report and asks whether the streaming platform should move. Nine months later the same executive is staring at a bill three times the original estimate, an exhausted platform team, and a half-migrated estate that now has to be finished because walking backwards costs more than going forward. The question was not unreasonable. Treating "migrate" as the natural next step was.

Most guidance about Apache Kafka assumes the hard problem is deciding where to move. The harder problem is deciding whether to move at all. A migration is a project with a budget, a risk profile, and an opportunity cost, and it competes against every other project the team could run instead. For a meaningful share of estates, the right answer right now is to stay, write the reasons down, and spend the budget somewhere with a better return. What follows is a checklist for reaching that answer honestly, because a decision to stay that nobody can defend collapses the next time someone in a review asks the question again.

1The migration nobody needed

Migration has a seductive property: it converts an operating problem into a project with a beginning, middle, and end. An existing estate is full of slow, chronic problems. Consumer lag that never stabilizes, partition counts that make rebalances painful, storage bills that grow with retention, on-call pages that everyone has learned to auto-acknowledge. A migration promises to erase all of them at once, on a schedule, with a named owner. That promise is what gets the project approved, and it is also what makes the project underestimated.

The chronic problems usually do not move with the platform. They live in the workload, in the data contracts, and in the team's capacity, which is why platforms get blamed for patterns they did not create. A topic that accumulates hot partitions will accumulate them on the destination too. A consumer that commits offsets on a timer instead of after processing will lag the same way on a new cluster. The migration changes the label on the cluster; it does not change the behavior of the people and systems around it.

Migrations are measured against the problems they actually remove, not the problems they were promised to remove.

The clearest sign a migration is being chosen for the wrong reason is that nobody can state which problem it solves. When the answer is "everything," the project has no exit condition, and a project without an exit condition cannot budget, cannot roll back, and cannot be declared finished. A decision to stay starts by refusing that framing and naming the precise condition under which moving would be justified. If that condition cannot be named, nothing else in the plan is load-bearing.

2Three stop signals: contract, bandwidth, shape

A three-light signal board for Kafka migration decisions, with contract, team bandwidth, and data shape as the stop signals checked before any cutover plan

Not every estate earns a stay decision. The ones that do usually trip one of three stop signals, each of which is visible before a single topic is copied. The order matters: contract governs what is even allowed this quarter, bandwidth governs what the team can absorb, and shape governs whether the destination architecture actually fits the data. A team one year into a multi-year commitment, with a hiring freeze and a retirement-heavy consumer fleet, is often reading all three lights at once.

  • Contract. The current platform is under an active term, a committed spend, or a negotiated discount that does not transfer, and leaving early converts a sunk commitment into an additional exit cost paid on top of the move. The signal is not the existence of a contract. It is a contract whose remaining term is longer than the time the migration would take to plan and execute.
  • Bandwidth. A platform migration consumes the same engineers who currently keep the platform alive. If the team is already running hot on upgrades, compliance work, or reliability fixes, the move borrows against production stability, and the first casualty is the cutover rollback nobody had time to rehearse.
  • Shape. The destination has an architecture, and the estate has a shape: retention length, partition counts, fan-out patterns, peak-to-average ratio, and the mix of hot and cold reads. A workload with extreme storage growth and long retention may benefit from an object-storage-backed design; a workload of small topics with heavy per-broker compute and strict low-latency writes may be indifferent or worse off.

The practical rule these lights produce is simple to state and harder to follow: one red light argues for delay, two argue for a written stay decision, and three mean the migration proposal should not come back until one of the lights changes. Unlike contract and bandwidth, shape answers itself with data rather than with a calendar, which may be why teams habitually check it last.

3A quick payoff estimate that ends the argument

A payoff estimate card for a Kafka migration, laying out one-time cost, recurring delta, avoided cost, and the break-even gate with evidence columns

The way to settle a migration argument quickly is to stop arguing about the platform and estimate the payoff. The estimate needs four numbers, all of which can be produced from invoices, cloud pricing pages, and the team's own runbook timings. If a number cannot be produced, its honest value is "unknown," and "unknown" belongs in the table because an estimate built on unknowns is the clearest reason not to migrate on a deadline.

LineWhat it measuresTypical result worth checking
One-time costMigration tooling, dual-run infrastructure, engineer weeks, retrainingLargest single entry; often rivals a year of platform spend for large estates
Recurring deltaNew platform cost minus old platform cost at the same retention and trafficCan be negative, flat, or positive depending on data shape
Avoided costRenewal, support, and operational toil removed over the forecast horizonUsually positive but slower than the sales narrative implies
Break-even gateMonths until accumulated delta exceeds the one-time costIf longer than the contract decision window, the move loses money now

The table's value is the break-even gate, not the individual cells. A migration can have a negative recurring delta and still be wrong, because a one-time cost that takes three years to earn back may be interrupted by the next architectural shift, the next contract change, or the next reorg before it pays off. The estimate converts "should we move" into "when does moving pay for itself," which is a question with a number for an answer and a number that can be argued with evidence instead of enthusiasm.

Two cells deserve extra suspicion because they are where estimates go wrong. One-time cost is habitually underestimated by a factor of two or more, because teams count topic copying and forget connector offsets, schema migrations, client bootstrap changes, private network routes, and the dual-run period when two platforms are both live. Recurring delta is habitually overestimated as a saving because it is computed against list prices rather than the negotiated terms the team actually holds. Correcting these two cells is often enough to flip the gate from "go now" to "not yet," and no platform comparison was needed to get there.

Before the estimate runs, remember what it is. The table is a decision tool, not a forecast. Its purpose is to make the migration earn its budget by showing the payoff horizon, and it works the same way when the horizon is short. A migration that pays for itself in two quarters with a rehearsed rollback should proceed on its merits. The point of the estimate is that it ends up on the right side of the argument for the right reason.

4Not migrating does not mean not evaluating

A stay decision is not a permanent answer, and the teams that treat it as one are the teams that rediscover the problem years later with fewer options. Staying put should be paired with a reassessment trigger: a dated condition under which the question reopens with a shorter, sharper scope. The trigger is what makes the stay decision durable, because the platform team can tell the next executive review that the estate is not stuck, it is parked behind an explicit gate.

Reassessment triggers that have survived real reviews tend to be measurable and rare rather than calendar-driven:

  • The current contract reaches a real exit window where leaving no longer burns the remaining term.
  • Retention, traffic, or partition counts cross a threshold that changes the estate's dominant cost, and the shape argument can be re-scored with new data.
  • A specific architecture capability that solves a named problem becomes available and verifiable, so the evaluation narrows to "does this actually change our shape" instead of "should we move again."

That third trigger is where a candidate platform enters the picture, and it enters late, only when the stop signals have already cleared. The kinds of estates that flagged the shape signal often share one trait: broker-attached storage makes retention and scaling expensive in ways that compound. Apache Kafka's Tiered Storage proposal, KIP-405, already separated remote log data from local storage as the reference for judging that class of design. If the estate's dominant cost is storage and data movement, a platform that removes broker-local persistence is a candidate worth scoring; if the dominant cost is elsewhere, the same platform is not a reason to reopen anything.

That is the honest boundary for AutoMQ. AutoMQ is a Kafka-compatible cloud-native streaming platform that separates compute from storage by writing stream data to object storage, which changes the economics of retention-heavy, cross-AZ estates. None of that makes it the right destination for a migration that should stay put. It makes it a reference to consult when the shape trigger fires, because the compatibility documentation and the architecture overview let the team test the shape hypothesis directly. Move only when the benefit of a Shared Storage architecture is confirmed against the estate's own retention and traffic data, not because the option exists.

5Decision states: migrate, stay, reassess

A decision states diagram for Kafka platform changes, showing migrate, stay, and reassess as explicit outcomes with the gates that move between them

The argument closes with a decision state, and it helps to treat the three outcomes as peers rather than as one right answer and two failures. "Stay" is a decision with its own evidence and its own owner. "Reassess" is a decision with a trigger date. "Migrate" is a decision with a plan. Each one is only as good as the written record that produced it.

StateWhat was decidedWhat must be written down
MigrateThe payoff estimate clears the break-even gate and the stop signals are greenScope, budget, exit condition, rehearsed rollback
StayOne or more stop signals are red and the estimate does not clearThe active signal, the rejected alternative, the date to re-score
ReassessThe shape argument is plausible but unproven against current dataThe trigger condition and the evidence that would fire it

The difference between a stay decision that holds and one that gets relitigated every quarter is the written record. A stay without a trigger is avoidance, and avoidance reads to the next reviewer as a failure to decide. A stay with a trigger reads as what it is: a platform team that priced the move, found it did not clear the gate, and recorded the condition under which it would.

Notice what is missing from all three states. None of them requires the team to prove the current platform is good, only that the alternative has not cleared its own bar. That reframing removes the debt that most stalled migrations begin with, which is the belief that staying means defending every wart on the current estate. Staying means the plan to leave was not good enough yet.

Back to the room where this started. The executive asked whether the platform should move, and the honest answer was that the question was premature. The contract had a year left, the team was fully assigned, and the data shape did not match the destination's mechanism tightly enough to pay for the move. The team wrote down all three lights, set a trigger on the contract window and the retention forecast, and spent the would-be migration budget on the consumer lag that had been annoying everyone for quarters. That is the quiet victory of a stay decision done well: not the migration that was avoided, but the expensive detour that never needed to happen. If the estate still needs to move later, evaluate the AutoMQ Open Source project on its own terms, and let the three lights decide.

6References

7FAQ

7.1When is not migrating Kafka the right call?

When one of three stop signals is red and the payoff estimate does not clear the break-even gate. An active contract with a remaining term longer than the migration would take is the first signal, because leaving early adds an exit cost on top of the move. A team with no spare capacity is the second, because a migration borrows against production stability. A data shape that does not match the destination architecture's dominant cost mechanism is the third. Staying is a decision, not a default, and it should be written down with the same rigor as a go decision.

7.2How do you estimate whether a Kafka migration is worth it?

Produce four numbers and gate on the break-even point: the one-time cost, the recurring delta between old and new platforms at the same retention and traffic, the avoided cost, and the number of months until the delta pays back the one-time cost. Correct the two cells where estimates break. One-time cost is usually undercounted because connector offsets, schemas, bootstrap changes, and the dual-run period are left out. Recurring delta is usually overcounted because list prices are compared instead of negotiated terms.

7.3What is a reassessment trigger for a stay decision?

A dated, measurable condition that reopens the migration question with a narrow scope. Useful triggers include the contract reaching a real exit window, retention or traffic crossing a threshold that changes the estate's dominant cost, and a specific architecture capability becoming available and verifiable. The trigger is what separates a stay decision from avoidance, because the next reviewer can see the condition that would reopen the question instead of relitigating it from scratch.

7.4Does evaluating a Shared Storage platform mean it is time to migrate?

It means the shape signal may have changed, not that migration should start. A platform like AutoMQ separates compute from storage by writing stream data to object storage, which changes the economics of retention-heavy, cross-AZ estates. That deserves scoring when the estate's dominant cost is storage and data movement. It does not deserve scoring when the dominant cost is elsewhere. The same payoff estimate applies, and the move happens only when the benefit of a Shared Storage architecture is confirmed against the estate's own data.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.