Blog

Migrating Kafka to GCP: Offset, ACL, and Rollback Gates

Table of Contents

Table of Contents

Moving Apache Kafka data to Google Cloud is straightforward to describe and difficult to finish. A replication connector can show a healthy task while consumers still point at the wrong position, producers authenticate with an identity that has no topic access, or the team has no safe way to return traffic to the source cluster.

The migration is complete when the target can serve the same workload contract: the required records exist, consumer groups resume at an understood position, principals have the intended permissions, and operators can stop or reverse the change without creating two writers. Treat a Kafka migration to GCP as a sequence of evidence gates, not as a single cutover command.

Migration state machine from inventory through replication, cutover, and rollback

1Start with a migration contract

Before deploying MirrorMaker 2.0, write down what “same workload” means for each topic and consumer group. Google Cloud’s Kafka migration guide uses Kafka Connect with MirrorMaker 2.0 and calls out topic, partition, throughput, producer, consumer, and consumer-group inventory.

For every workload, capture the source fact, the target expectation, and the artifact that will prove the two agree:

Contract areaTarget expectationEvidence to keep
RecordsThe selected topics and partitions contain the required history through the cut lineProducer event IDs, partition counts, and replication checkpoints
ProgressEach consumer group resumes from a documented translated offset or timestamp policyGroup offsets before cutover, checkpoint records, and first target commits
IdentityProducers, consumers, connectors, and operators use the intended target principalsIAM bindings, client configuration, and authorization test results
OwnershipOne cluster is the write authority at every point in the changeCutover decision, traffic switch record, and source write status
RecoveryThe previous path can be restored without split-brain writesRollback trigger, endpoint change, and post-rollback consumer evidence

The cut line is the reference for every comparison. Mark the last source event that the migration accepts, then use event IDs and partition positions to explain anything seen on the target.

2Replicate records, checkpoints, and ACLs as separate paths

MirrorMaker 2.0 has different connectors for different state. Google’s migration documentation describes a Source connector for topic data and a Checkpoint connector for consumer checkpoints. The Source connector can create target topics and replicate topic ACLs, while the Checkpoint connector copies consumer offsets and translates them into the target cluster’s offset space. These paths must be monitored and tested independently; a green source task does not prove that a group can resume.

The source and target also have different control boundaries. A target cluster may use a new project, Virtual Private Cloud (VPC), service account, or authentication mechanism even when the Kafka client code stays the same. Keep the connector’s replication identity separate from application identities, and record which identity is used for discovery, writes, checkpoint synchronization, and administration.

Google Cloud’s Kafka ACL documentation describes ACLs as fine-grained authorization within the Kafka cluster. It also explains that the principal used by the standard Kafka authorizer is derived from the authenticated identity; for SASL, that can be an Identity and Access Management (IAM) principal such as a service-account email. A source ACL whose principal is a local username will not automatically become the intended Google Cloud principal. Define the mapping explicitly and test it with the real client protocol.

Offset translation and ACL identity mapping from a source Kafka cluster to GCP

Keep these paths visible in the migration runbook:

  • Record path: source producer → Source connector → target topic and partition.
  • Progress path: source consumer group → Checkpoint connector → translated target offset.
  • Authorization path: client identity → target authentication → target ACL evaluation.
  • Control path: connector and operator identity → topic, connector, and cluster APIs.

They can fail in different ways: records may be current while checkpoints lag, or checkpoints may be present while a target principal is denied.

3Build an offset proof instead of restarting consumers

Offsets are positions within a partition, not global event identifiers. MirrorMaker 2.0 can translate checkpoints, but the precision of that translation depends on the connector settings and the relationship between source and target records. Google’s migration guide documents settings such as offset.lag.max, sync.group.offsets.enabled, and sync.group.offsets.interval.seconds; verify the values used by your deployment rather than treating the defaults as a guarantee.

The offset gate should answer three questions for every migrated group:

  1. What was the source position? Capture the committed offset, assignment, and latest source position at the cut line.
  2. What is the target position? Record the translated checkpoint and the target partition where it will be applied.
  3. What did the application observe? Consume a bounded sample, compare producer-generated event IDs, and record duplicates, gaps, ordering, and sink results.

Run the test with the same consumer configuration that will serve production traffic. A fresh group that reads from the beginning proves that the target has bytes, but not that an existing group can resume. If the workload uses a timestamp fallback, record the timestamp and resulting event-ID range so that a later duplicate or gap can be explained.

Offset synchronization is also time-sensitive. If the checkpoint task is behind the source, a consumer cut over too early can replay an older position. If it is stopped too late, the target may continue receiving data after the team has declared the source authoritative. Set a measurable checkpoint freshness condition, pause the cutover when it is not met, and keep the artifact with the change record.

4Recreate ACLs as an identity contract

Permissions are part of the application interface. A producer that reaches the target but receives TopicAuthorizationException is not migrated, even if every record is present. The target ACL review should begin with the client inventory, then map each source principal to a target identity and a minimal set of operations on topics, groups, transactional IDs, and cluster resources.

Google Cloud’s ACL model retains Apache Kafka’s resource and operation vocabulary, while the authenticated principal can come from IAM. The service documentation also describes behavior when no ACL is defined for a resource. Do not rely on that behavior as an accidental migration shortcut: export the intended ACL set, apply it to the target, and test both allowed and denied operations before redirecting traffic.

Client roleRequired target testFailure that blocks cutover
ProducerAuthenticate, describe the topic if required, and produce a test recordAuthentication or write authorization error
ConsumerJoin the group, fetch the assigned partitions, and commit a test offsetGroup or read authorization error
ConnectorDiscover topics, create or update target resources, and replicate checkpointsMissing administrative or data permissions
OperatorList and inspect topics, groups, and connector status through the approved pathRunbook depends on an unapproved identity

Store the target ACL export and denied-operation test output with the migration evidence. A successful producer test cannot prove that every consumer group or connector path is authorized.

5Put cutover behind explicit gates

The cutover should be a short, reversible change because the long work has already happened in the gates. Use a decision record with named owners and stop conditions:

  1. Inventory gate. Topics, partitions, retention, producers, consumers, groups, connectors, and source ACLs are complete.
  2. Replication gate. Source and checkpoint connectors are healthy, target topics match the intended shape, and the cut line is recorded.
  3. Offset gate. Every in-scope group has a translated position or an approved timestamp policy, with event-ID evidence from a target replay.
  4. ACL gate. Target principals pass allowed-operation tests and fail the denied-operation tests that protect unrelated topics and groups.
  5. Traffic gate. The target producer and consumer paths are reachable from the production network, and the switch mechanism has been rehearsed.
  6. Rollback gate. The team can identify the trigger, stop target writes, redirect clients, and verify that source consumers resume without a second writer.

Migration gates for inventory, replication, offset, ACL, traffic, and rollback

Do not compress these into a single “replication healthy” check. Each gate protects a different failure mode, and each should have an artifact that an on-call engineer can inspect without reconstructing the migration from chat messages.

6Rollback is a write-path decision

Rollback becomes unsafe when both clusters accept writes. Decide in advance how the application handles a failed target canary: stop target producers, fence the target endpoint, and redirect clients to the source only after the source’s last authoritative position is known. If the target has accepted records that do not exist on the source, the owner must choose whether to replay them, reconcile them, or declare them outside the source contract. There is no universal “reverse” operation that preserves business meaning for every workload.

Use an observable rollback trigger, such as authorization failures across an in-scope client class, a checkpoint freshness breach, or a record-ID mismatch in the canary sink. Capture the target’s last accepted event ID and consumer positions before switching back. After rollback, verify source writes, source consumer position, and that the target is no longer a writer. Keep the target available for investigation only after it has been fenced from production writes.

Run the sequence as a rehearsal with a bounded topic and representative group, then inspect the artifacts as if the change had failed.

7Where Kafka-compatible shared storage fits

The gate framework applies whether the destination is a Google-managed Kafka service, Kafka on GKE, or another Kafka-compatible deployment. Storage architecture changes which broker replacement and recovery assumptions need testing; in a broker-local design, moving the cluster can also move retained bytes and replica state.

AutoMQ is a Kafka-compatible cloud-native streaming platform built on a Shared Storage architecture. Its S3Stream layer places durable stream data in object storage while brokers provide Kafka protocol handling and compute. That gives a migration evaluation a separate storage boundary to test while Kafka clients, topics, offsets, and ACLs remain in the contract. MirrorMaker behavior, principal mapping, checkpoint freshness, and rollback fencing still need evidence.

For an AutoMQ proof, reuse the same source event IDs, consumer groups, ACL matrix, and cutover gates. Test a broker replacement or staged deployment according to the target runbook, then compare record integrity, offset behavior, authorization, and rollback evidence with the GCP baseline. The useful result is a measured migration boundary, not a promise that one storage model fits every workload.

8FAQ

8.1Does MirrorMaker 2.0 migrate consumer offsets automatically?

It can replicate and translate consumer checkpoints through the Checkpoint connector, but the result depends on connector configuration, source and target positions, and checkpoint freshness. Test the translated target position with event IDs before cutover.

8.2Do Kafka ACLs copy cleanly to GCP?

MirrorMaker 2.0’s Source connector can replicate topic ACLs, but the target still evaluates the authenticated target principal. Map local principals to the intended IAM or certificate identity and run allowed and denied-operation tests.

8.3What is the safest rollback point?

The safest point is before the target becomes a second writer. If a target canary has already accepted records, capture its last event and consumer positions, fence target writes, and apply an application-specific reconciliation policy before returning to the source.

8.4Can I validate a migration with a new consumer group?

A new group proves that records can be read. It does not prove that production groups resume from the intended position. Include a bounded replay with a translated checkpoint or a documented timestamp policy.

The migration is ready when a reviewer can trace one producer event from the source, through replication, into a target partition, past an authenticated client, and back through a tested rollback path. Keep the offset proof and ACL export beside the cutover decision. To evaluate the same Kafka-compatible contracts on a shared-storage deployment, start with AutoMQ Open Source.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.