Blog

Redpanda Multi-Region Kafka DR: RPO, RTO, and Failback Evidence

Table of Contents

Table of Contents

A disaster-recovery diagram can pass review quickly. A recovery test is harder because Kafka recovery is a chain of state transitions: records must cross a region boundary, a target must become writable, producers must reconnect, consumer groups must resume from a defensible position, and downstream systems must accept the result. If one link is missing, a green cluster-health dashboard can still hide data loss or an unbounded recovery window.

For Redpanda, the documentation accessed for this draft describes Shadowing as asynchronous, offset-preserving replication from a source cluster to a Redpanda shadow cluster. The shadow cluster is read-only until a failover, and the documented model is active-passive rather than active-active. Those are useful architectural facts, but they are not an RPO or RTO promise for your workload. RPO is determined by the replication boundary at the moment of failure; RTO is determined by the complete failover procedure and the applications that depend on it.

That distinction gives a Kafka platform team a practical test: replace “multi-region DR is supported” with an evidence pack that says which records were protected, how long the service took to recover, and what happened when the original region returned.

mermaid
flowchart LR P[Source Redpanda cluster\nproduction writes] -->|asynchronous Shadowing| S[Shadow cluster\nread-only replica] S -->|failover after evidence gate| W[Promoted writable cluster] W --> C[Kafka producers and consumers] C --> V[Business validation\nevent IDs, lag, outcomes]

1RPO and RTO are workload measurements

Recovery Point Objective (RPO) answers a record question: which acknowledged data may be absent from the recovery target when the source fails? For Kafka, a timestamp alone is not enough. The same wall-clock interval can contain very different record volumes across partitions, and a producer retry can make a timestamp comparison look healthy while the application sees a duplicate or a gap. Measure the boundary with partition positions and stable application event identifiers, then describe the result in the terms your workload actually uses.

Recovery Time Objective (RTO) answers a service question: how long from the declared incident until the workload is producing, consuming, and passing its business acceptance checks on the recovery cluster? A broker becoming reachable is only an intermediate event. Client DNS, TLS or SASL credentials, topic write permissions, Consumer group offsets, schema services, connectors, and downstream sinks can all extend the real recovery window.

Capture the following evidence for each DR exercise:

EvidenceWhat it establishesWhat it cannot prove by itself
Per-partition replication lagHow far the shadow is behind the source at the failure boundaryThat every acknowledged record is readable by the application
Last source event and first target eventThe record interval that must be reconciledThat consumers resumed at the intended position
Producer acknowledgment timelineWhen writes succeeded on the promoted clusterThat stale producers are unable to write to the old region
Consumer assignment and processing timelineWhen consumers joined, received partitions, and made progressThat downstream side effects were applied exactly once
Business reconciliationWhether event IDs, counts, and outcomes match the agreed policyThat the next failover will have the same result without another drill

Do not publish a single RPO or RTO number from a vendor page. Set an objective for a named workload, record the topology and client versions, and report the measured result with the test date and failure scenario.

2What Redpanda Shadowing changes in the DR model

Redpanda’s Shadowing overview states that Shadowing uses asynchronous replication and keeps the source cluster active while the shadow cluster receives updates in read-only mode. The source can be another Redpanda cluster or a Kafka API-compatible source, but the target-side behavior still depends on the version, license scope, topic filters, security configuration, and workload. Treat those as test inputs rather than assumptions about every Redpanda deployment.

The same documentation describes replication of topic data, topic properties, Consumer group offsets, and cluster metadata. It also says that Shadowing operates asynchronously, so replication lag exists. That lag is the starting point for an RPO test, not a failure of the design. The correct question is whether the observed lag remains inside the business tolerance during normal operation and whether the team knows what to do when it grows.

The failover documentation adds boundaries that belong in the runbook. It says in-flight transactions at the source are not replicated and may be lost, and that writing to both clusters after a network partition can create metadata divergence. It also states that automatic fallback to the original source cluster is not supported after failover. These points make failback a separate operation with its own writer-fencing, reconciliation, and client-redirection gates.

The architecture can be summarized as a state machine:

mermaid
stateDiagram-v2 [*] --> ActiveSource ActiveSource --> ShadowingHealthy: link active, lag within objective ShadowingHealthy --> FailoverDeclared: source region unavailable FailoverDeclared --> TargetPromoted: topics promoted and clients redirected TargetPromoted --> FailbackPreparation: source region restored FailbackPreparation --> ActiveSource: reverse path reconciled and one writer selected TargetPromoted --> IncidentHold: dual writes or unexplained gaps IncidentHold --> FailbackPreparation: owner and evidence gate cleared

A team that cannot name the owner of each transition does not yet have a DR runbook. “The platform will fail over” is not an owner; the runbook needs a person or automation boundary for declaring the incident, fencing writers, promoting topics, changing client endpoints, and approving failback.

3Build the pre-failover evidence pack

Start with an inventory that can be replayed during the exercise. Record source and shadow cluster identifiers, Redpanda release, deployment substrate, licensing or feature scope, link configuration, authentication method, topic filters, partition counts, retention policies, and the consumer groups that matter to the business. Include Schema Registry, Kafka Connect, Kafka Streams, Flink, REST clients, DNS, secrets, and sink systems where they participate in the data path. A replicated topic does not automatically mean every surrounding state has been replicated. A Kafka DR evidence worksheet can help keep the recovery boundary and data-export assumptions visible during review.

Before injecting a failure, take a baseline from the source and shadow. Redpanda’s monitoring guidance documents the redpanda_shadow_link_shadow_lag metric, which compares the source partition’s Last Stable Offset with the shadow partition’s High Watermark. Capture that value by topic and partition together with client production rate, consumer lag, shadow link state, topic state, and client errors. Keep the raw time window; a single screenshot cannot show whether lag is stable, growing, or recovering.

Use a pass gate that matches the workload, for example:

  • The shadow link is healthy and every in-scope topic is in the expected state.
  • Replication lag is measured against the workload’s RPO objective, with no unexplained growth.
  • The target has the required topic configuration, credentials, schema subjects, and Consumer group state.
  • Producers and consumers have a tested endpoint change, and the old endpoint can be fenced.
  • The business owner has defined the event identity, duplicate policy, and rollback decision.

A pass gate is deliberately more demanding than “the link is active.” If the test does not capture client and business dependencies, it cannot explain a failed recovery later.

4Measure RPO at the record boundary

When the source fails, freeze the evidence before attempting to repair it. Record the source’s last reachable partition positions, the shadow’s corresponding positions, and the timestamps of the last acknowledged source writes. If the workload carries a stable event ID, sample records around the boundary and compare event IDs, keys, timestamps, and partition assignments on the shadow. If it does not carry a stable ID, say so in the result; the absence of an identity key limits what the exercise can prove.

The RPO result should answer four questions:

  1. Which source records were acknowledged before the failure?
  2. Which of those records are present and readable on the shadow?
  3. Which records were duplicated during retries or replay?
  4. Which records cannot be reconciled because the application has no stable identity?

Do not turn a lag gauge into a promise such as “zero data loss.” Asynchronous replication can be healthy while still leaving a non-zero recovery boundary at the instant of failure. Conversely, a short lag interval does not prove that a transaction committed at the source is visible on the target. Redpanda documents the special case explicitly: source in-flight transactions are not replicated. Test committed records with read_committed consumers and account for aborted or unfinished transactions separately.

For high-value streams, make the reconciliation artifact part of the change record. Include the source and target offsets, event IDs, counts by partition, transaction status where relevant, and the decision about each missing or repeated record. The artifact is more useful than a headline RPO because it tells the next operator what “protected” meant for that workload.

5Measure RTO across the whole client path

Use one clock that starts when the incident is declared and ends only after the workload passes acceptance on the promoted cluster. Split the interval into observable milestones so the team can see where time was spent:

MilestoneEvidence to retain
Incident declaredIncident ID, failure scenario, and decision owner
Source writes fencedDeployment state, authorization result, and final source acknowledgment
Shadow topics promotedCommand or API response, topic state, and promotion timestamp
First target producer acknowledgmentClient logs and an application event ID
Consumer group resumedJoin, assignment, committed offset, processing rate, and lag
Business acceptance passedReconciliation output and downstream success signal

The Redpanda failover documentation recommends testing failover in a non-production environment to measure RTO. Use that advice as an engineering requirement, not a ceremonial exercise. A drill that measures only the administrative promotion command leaves DNS propagation, secret rotation, consumer rebalances, and sink recovery outside the result.

Failure injection should match the outage you care about. A blocked inter-region path tests a different boundary from a lost source cluster, a control-plane outage, or an authentication failure. Name the failure mode in the report and repeat the exercise after topology or client changes. RTO is a property of the runbook and the workload, not a permanent attribute of the product name.

6Treat failback as a separate recovery event

Failback is where many DR plans become vague. When the source region returns, do not point clients back only because its brokers are healthy. Redpanda documents that automatic fallback is not supported after Shadowing failover and warns against writing to both clusters. That means the return path needs a deliberate single-writer decision, data reconciliation, and an explicit client sequence.

A safe failback rehearsal can be organized around these gates:

  • Freeze the promoted cluster’s writers. Stop or fence producers before changing ownership. Record the last accepted target event and prevent stale clients from reconnecting through cached DNS or old credentials.
  • Reconcile the return boundary. Compare target records with the restored source by event ID, partition, timestamp, and business result. Decide how to handle records that were produced after the original source failed.
  • Prepare the source as the sole writer. Re-establish replication or the selected data-transfer path, verify topic and offset state, and keep the source in a non-writable recovery state until the owner approves the cutback.
  • Move consumers deliberately. Confirm the actual offset source for each consumer. A Kafka group offset, a Flink checkpoint, and an application-managed seek position are different recovery inputs.
  • Prove the return. Run a canary producer and consumer, validate event identity and business outcomes, then expand traffic according to the same acceptance gates used during failover.

The order matters because a numeric offset is local to a Kafka log. A client that reads the same number from a different cluster does not necessarily read the same record. If the application uses an external checkpoint or database, copying Kafka Consumer group offsets does not rewrite that state. Keep the offset source visible in the runbook and test it explicitly.

7Keep the platform comparison honest

The DR evidence should stand on its own before you compare vendors. Redpanda’s Shadowing documentation is a useful example of a design with clear boundaries: asynchronous replication, active-passive operation, read-only shadow topics before promotion, measurable lag, and a failback procedure that must be planned. The same questions apply to Apache Kafka with MirrorMaker, managed Kafka services, and other Kafka-compatible platforms. Ask what is replicated, which offsets are preserved, how transactions behave, who fences writers, and what happens when the original region returns.

AutoMQ can enter the discussion after those requirements are written down. AutoMQ is a Kafka-compatible cloud-native streaming platform with a Shared Storage architecture; its brokers use Kafka protocol semantics while durable stream data is held through S3Stream and object storage. That architecture is relevant when the DR review exposes a separate problem: broker-local storage makes recovery, scaling, or migration too tightly coupled to individual nodes. It does not turn asynchronous cross-region replication into a zero-RPO guarantee, and it does not remove application-level ownership or failback work.

For a migration or controlled platform change, AutoMQ commercial editions document a Kafka Linking guide. The guide covers source endpoints, replication lag, consumer offset sources, promotion gates, and rollback boundaries. It describes byte-to-byte synchronization and offset handling for the migration path; those claims still need to be validated against the named source version and the workload’s external state. A migration tool is not a substitute for a regional DR design, but its evidence discipline is useful when a team evaluates a candidate Kafka-compatible data plane. Teams comparing return paths can also use this failback validation runbook as a checklist for writer fencing, offset evidence, and client recovery.

8The runbook is the product

A multi-region Kafka DR design is ready when another operator can reproduce the result without asking what “caught up” means. Keep the topology, version, link state, lag series, writer-fence evidence, client timeline, offset source, event reconciliation, and failback decision together. Store the result with the change or incident record, and repeat the drill after changes to partitions, retention, authentication, schemas, or client libraries.

Return to the original question: how much data can this Redpanda Kafka workload lose, and how long until it serves correct traffic again? The answer belongs in the evidence pack, with a named failure scenario and a date. If the team cannot show the records and timestamps behind the answer, it has a recovery intention rather than a recovery result.

If you are evaluating a Kafka-compatible architecture after running this drill, bring the evidence to AutoMQ. Share one representative workload, its offset source, and the failover timeline; the useful conversation starts with the test boundary rather than a headline recovery number.

9FAQ

9.1Does Redpanda Shadowing guarantee a fixed RPO?

No fixed RPO should be inferred from the feature name or documentation. Shadowing is asynchronous, so replication lag exists. Measure per-partition lag and reconcile the records acknowledged before the failure against the shadow’s readable data for the named workload.

9.2Is Redpanda Shadowing active-active?

The Shadowing documentation describes an active-passive pattern: the source handles production traffic and the shadow cluster remains read-only until failover. Do not design a multi-writer topology unless the exact product and release documentation, licensing, and test evidence support it.

9.3What happens to in-flight Kafka transactions during failover?

Redpanda’s failover documentation says in-flight transactions at the source are not replicated. Test committed and aborted transactions separately, and use read_committed consumers plus application event IDs to reconcile the boundary.

9.4Can I automatically fail back when the source region returns?

Do not assume automatic fallback. Redpanda’s failover documentation states that automatic fallback to the original source is not supported after failover and warns against writing to both clusters. Plan failback as a separate single-writer event with reconciliation and client cutover gates.

9.5Does AutoMQ solve multi-region Kafka DR by itself?

No. AutoMQ’s Shared Storage architecture addresses Kafka compute and storage coupling, while Kafka Linking provides a documented migration path with evidence gates. Regional replication, writer ownership, RPO, RTO, and failback still need a workload-specific design and test.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.