Table of Contents
Table of Contents
An auditor asks for the Apache Kafka® exit plan. The platform team opens a folder, finds a diagram of two clusters, and then has to explain who can move the data, where consumers resume, and how long recovery is supposed to take. The diagram is not the plan. The plan is a set of measured recovery targets, an export path, a cutover boundary, and evidence that a named team has exercised all four.
RTO and RPO make the question testable. Recovery time objective, or RTO, is the elapsed time from declaring the recovery event to restoring the agreed business path. Recovery point objective, or RPO, is the age of the newest data point the target can prove it has recovered at that decision point. In Kafka, both values depend on record identity, partition position, consumer state, schemas, credentials, and the path used to move data.
Define the target per workload, then choose the export path that can produce evidence for it. A Kafka endpoint may reduce application changes, but it does not prove that records, consumer positions, or rollback state are recoverable.
1The auditor asks for a plan you can run
Start with one workload and one failure boundary. Write down the source cluster, candidate target, topics, important producers and consumer groups, retained history, and the downstream effect that counts as service restored. Keep one source of truth for the recovery declaration.
The plan should record four timestamps for every rehearsal:
- Incident time: when the data loss, service loss, or vendor exit condition is defined to have started.
- Recovery point: the last record or partition position that the target can verify as present and decodable.
- Recovery declaration: when the incident commander authorizes the target path to serve the workload.
- Service restored: when the agreed producer and consumer path passes its acceptance checks.
The timestamps turn vague promises into calculations. Observed RPO = incident time - recovery point time, measured as data age at the failure boundary. Observed RTO = service restored time - recovery declaration time. Write the workload's threshold as a variable until the team measures the path. A drill outside that threshold has found a design gap.
Kafka adds a boundary that generic disaster recovery templates often miss: a record can be present while the application cannot use it. The target must have the topic, partition, key, value, headers, timestamps, and serializer or schema needed to decode it. A consumer also needs a valid handoff position. A numeric offset is a local coordinate, so test offset translation or timestamp-based replay explicitly.
2RTO and RPO belong to the workload, not the contract template
Two topics on the same cluster can need different recovery paths. A payment authorization stream may require a narrow loss window and a carefully reconciled consumer handoff. A telemetry stream may be rebuilt from another source. An analytics stream may tolerate replay from a retained window. The Kafka endpoint is shared; the recovery contract is not.
Define these fields before evaluating tools:
| Field | Kafka-specific question | Evidence to retain |
|---|---|---|
| Recovery point | Which record, partition position, or application checkpoint is the latest acceptable target state? | Export lag, checkpoint, record identity, and decode result |
| Recovery time | What producer, consumer, connector, schema, and network checks must pass before service is restored? | Timestamped runbook events and acceptance output |
| Replay boundary | Which records may be processed again, and how are duplicates detected? | Idempotency key, reconciliation query, or downstream ledger |
| Rollback boundary | Which system remains authoritative while the target is being accepted? | Authority flag, freeze rule, and reverse path |
3Three export paths and the boundaries they create
Choose an export path whose failure modes match the RPO, RTO, and replay contract. Use more than one path when live cutover and historical backfill have different requirements.
| Export path | What it can preserve | Where the plan needs proof | Best fit |
|---|---|---|---|
| MirrorMaker 2 or another Kafka-aware replication path | Kafka records and selected topic metadata through a continuously operated bridge | Replication lag, topic naming, offset or checkpoint translation, schema handling, and reverse direction | A target that must stay close to the source while both clusters are available |
| Application dual write | Records at the application boundary, with the possibility of very small write divergence when both paths accept the same event | Idempotency, ordering, partial failure, retry behavior, backpressure, and a single authority for acknowledgments | Workloads whose producers can own duplicate detection and controlled cutover |
| Object storage export or snapshot | A durable export or snapshot when the platform documents a supported restore contract and the format is usable by the target | Storage layout, metadata, encryption, versioning, restore tooling, record decoding, and point-in-time consistency | Historical backfill or platforms whose durable data path has a documented portable snapshot interface |
MirrorMaker 2 works at Kafka's record boundary and can run while the source remains authoritative. A mirrored topic, a translated offset, and a schema registry subject are separate recovery objects. Treat checkpoints as evidence to verify rather than proof that every consumer resumes correctly.
Dual write moves the problem closer to the producer. It can give the application a precise event identity and a place to reconcile duplicates, while making the application responsible for partial failure, retries, and ordering. Use it only when the team can show how acknowledgments, idempotency, and rollback behave.
An object storage snapshot can help with history, but files in a bucket are not automatically a portable Kafka export. Broker log segments, indexes, checkpoints, and metadata may depend on the source implementation. Use this path only when the source documents a restore contract, the target can consume the result, and a rehearsal restores a sample through Kafka semantics.
For a platform built around object storage, the storage boundary is a fact to investigate rather than a shortcut to assume. AutoMQ is a Kafka-compatible cloud-native streaming platform with durable stream data in object storage. Its Kafka compatibility documentation and architecture documentation identify boundaries to test.
That combination affects an export and cutback review in two ways. Kafka compatibility may reduce client changes, while object storage may offer a different place to inspect retention and restore boundaries. Neither fact guarantees that a bucket can attach to another Kafka system. Verify ownership, metadata, encryption, versioning, restore procedure, consumer checkpoints, and target acceptance tests.
4Cutover is a point, and cutback is a separate plan
An exit plan needs a cutover point that an operator can identify during an incident. “The target is caught up” is too vague. Define the point with a source partition position, a timestamp plus record identity, or an application checkpoint. State how the target proves it received and decoded records produced during the final drain.
Keep the authority boundary visible during the transition. A common sequence is to keep the source authoritative, pause or fence writes according to the workload contract, drain the export path, validate the target checkpoint, start target consumers with an explicit replay policy, and then change producer routing. The actual order depends on the export mechanism. The plan must show which step changes the authority and which step can still be reversed.
Cutback deserves its own acceptance criteria. If the target accepted records after cutover, returning to the source requires a reverse export, a reconciliation window, or a decision to discard target-only records. A DNS change moves clients but does not reconcile two histories. Name the last source-authoritative point, target-only data handling, consumer state, schema changes, connector positions, and the person who can declare the return safe.
See the managed Kafka disaster recovery guide for failure domains and failover ownership. This article focuses on the reversible record boundary that a second cluster does not prove.
5A rehearsal calendar with named owners
Evidence expires when clients, schemas, connectors, networks, or retention policies change. Combine recurring checks with change-triggered drills, and assign an evidence owner, a decision approver, and a runbook operator.
| Cadence | Exercise | Primary owner | Pass evidence |
|---|---|---|---|
| Continuous | Record export lag, last verified checkpoint, target decode errors, and source authority | Platform SRE | Dashboard or log link tied to the workload and time window |
| Monthly | Restore or replay a representative topic window and validate keys, headers, schemas, and compaction behavior where relevant | Data platform owner | Reconciliation result with sample identities and an exception owner |
| Quarterly | Run a controlled cutover, consumer handoff, downstream validation, and cutback | Incident commander with application owner | Timestamped RPO/RTO calculation, decision record, and rollback result |
| After a material change | Repeat the affected path after a client, schema, connector, network, retention, or platform change | Change owner | Change record linked to updated runbook evidence |
| Before renewal or exit | Review all open gaps, target access, export permissions, credentials, and contract dependencies | CTO or platform director | Signed inventory with owners, expiry dates, and go or no-go decision |
The drill should be repeatable and real enough to expose failure. Select a representative topic, producer path, consumer group, relevant connector, schema versions, and cutover network identity. Keep the source authoritative during early exercises. Test target writes separately for duplicate and reverse-path behavior.
The cross-region Kafka recovery drill guide and dual-write cutover review cover adjacent drill questions. Keep your workload evidence as the decision record.
6The audit-ready exit plan template
Store one record per workload. A useful entry can fit on a page, but it should answer every question an operator will face:
Workload and business owner:
Source cluster and target:
Topics, partitions, retention, and schemas in scope:
RPO target and measurement definition:
RTO target and service-restored definition:
Export path and historical backfill path:
Recovery point identifier:
Consumer handoff and replay policy:
Source authority boundary:
Cutover procedure:
Cutback procedure and target-only data handling:
Platform SRE, application owner, data owner, network owner:
Evidence links and expiry dates:
Open gaps, approver, and next rehearsal:
This template makes gaps visible. Mark untested fields as gaps and schedule them before the exit is needed.
The auditor who asks for an exit plan is really asking whether the organization can prove control of its data under pressure. Give that question a recovery point, a recovery time, a named export path, and a person who has run the cutback. To test the same boundaries against a Kafka-compatible target, start an AutoMQ evaluation with one representative workload and the acceptance fields above.
7References
- Apache Kafka documentation, including cross-cluster replication and MirrorMaker 2 concepts.
- Apache Kafka replication design, for partition replicas, leaders, and in-sync replicas.
- Apache Kafka Connect documentation, for the connector-based replication boundary.
- Apache Kafka topic configuration, for retention and cleanup behavior that affects export and replay.
- AutoMQ compatibility with Apache Kafka.
- AutoMQ architecture overview.
- AutoMQ S3Stream shared streaming storage.
8FAQ
8.1What is an RPO in Kafka?
RPO is the data age the business accepts at the recovery boundary. Measure it from the defined incident time to the newest target record, partition position, or application checkpoint that the team can verify as present and decodable.
8.2What is an RTO in Kafka?
RTO is the elapsed time from recovery declaration to restoration of the agreed business path. Include target readiness, producer routing, consumer handoff, schema and credential checks, and downstream acceptance criteria.
8.3Is MirrorMaker 2 enough for a Kafka vendor exit plan?
It can be part of the live export path, but the plan still needs tests for topic naming, lag, consumer checkpoints, schemas, replay, and reverse direction. A replicated topic is one piece of a recoverable application path.
8.4Can Kafka broker files be copied to another vendor?
Do not assume broker log files are portable. They can depend on the source cluster's storage layout, indexes, checkpoints, and metadata. Plan around a supported Kafka-aware export, a documented snapshot restore contract, or an application-level replay path.
8.5How often should a Kafka exit plan be rehearsed?
Keep export evidence visible continuously, run representative restores on a recurring cadence, and repeat affected tests after material changes. Record the scope and measured RPO/RTO every time.
8.6Where does AutoMQ fit in a Kafka exit plan?
AutoMQ can be evaluated when Kafka protocol compatibility and object-storage-backed durable data are part of the target architecture. The evaluation still needs to verify record export, metadata, schemas, consumer checkpoints, restore behavior, and the cutback boundary for the workload.
