Table of Contents
Table of Contents
Terraform can create a Kafka cluster on Google Cloud in one apply. The harder question is whether a second plan is empty, whether an operator can explain every dependency in the graph, and whether a failed change can be reversed without guessing which cloud resource moved first.
For GCP Kafka, infrastructure as code is a lifecycle contract. It should describe the service boundary, identity, network path, secret handling, and evidence required after a plan or rollback.
1Define the desired state before the resource block
Start with a short contract that a reviewer can read without opening Terraform. The contract should answer what is being managed, where it lives, and what the team is deliberately leaving to a provider or a separate system. A managed Kafka cluster, its topic policy, and its client attachment may belong in one state boundary. A consumer deployment, application schema, or data migration may need a different owner and a different state file.
The Google Cloud Terraform documentation describes the provider workflow, while the Managed Service for Apache Kafka Terraform guide lists resources for clusters, topics, ACLs, Connect clusters, and connectors. Those pages are syntax references; they do not decide the ownership boundary for your organization. Write that boundary down first.
| State boundary | Put under Terraform | Keep as a separate contract |
|---|---|---|
| Platform foundation | Project, APIs, service identity, VPC attachment, private access | Organization policy and billing approval owned by a central team |
| Kafka service | Cluster configuration, approved locations, capacity shape, access controls | Workload-level SLOs and application deployment |
| Data objects | Topics, ACLs, connector declarations where the provider supports them | Message schemas, replay jobs, and data-quality checks |
| Recovery | Backup or export references, restore prerequisites, runbook links | The incident decision and human approval to restore |
Clear boundaries make a plan reviewable. If a resource is created by a pipeline but changed through a console, mark that exception instead of pretending it is fully declarative.
2Build a dependency graph that exposes failure order
Terraform's graph is more useful when its dependencies describe runtime needs. A Kafka cluster may depend on an enabled API, a service account, IAM bindings, a network path, and a region-specific configuration. Topics and ACLs depend on the cluster becoming reachable. Connector workers depend on Kafka access and the destination's identity. A graph that omits those relationships can produce a successful apply that fails when the first client connects.
Keep the graph legible by separating three kinds of dependency:
- Creation dependencies: a resource cannot be created until another resource exists, such as a cluster requiring a project API.
- Access dependencies: a resource exists, but a caller cannot use it until IAM, DNS, routes, or firewall rules are ready.
- Operational dependencies: a change is technically possible but should wait for a review, a maintenance window, or a recovery test.
Use explicit depends_on only when Terraform cannot infer a real dependency. Overusing it serializes unrelated work and hides the actual graph. Prefer references and module outputs for ordinary creation relationships, then reserve depends_on for API enablement, policy propagation, or another boundary that the provider cannot express directly.
Walk the first plan review from the outside in: project and identity, network, Kafka service, data objects, then clients. Ask what happens if each layer is delayed or partially applied.
3Keep identity and secrets out of the state boundary
Kafka provisioning usually needs two different kinds of sensitive input. The first is an identity that Terraform uses to call Google Cloud APIs. The second is a credential or certificate that Kafka clients use after the service exists. They have different owners and different rotation paths.
Google Cloud IAM defines who may administer resources, while Secret Manager provides a managed place for application secrets. Neither service makes a secret safe if it is interpolated into a Terraform argument that the provider records in state. Before adding a variable, check whether the value will appear in a plan, a state snapshot, a module output, or a CI log.
An identity review asks:
- Which service account runs the plan and which one runs apply?
- Which API permissions are required to create the cluster, topics, ACLs, or connectors?
- Which identity can read client secrets, and is that identity different from the provisioning identity?
- How are credentials rotated without forcing an unrelated cluster replacement?
- Where is state stored, who can read it, and how is its history retained?
Provider examples may show a password argument or sample variable. Treat that sample as syntax guidance, not as a secret-management policy. Pass references where the resource supports them; otherwise, make state exposure an explicit risk and choose a different integration boundary.
4Idempotence means the second plan explains itself
An idempotent workflow converges on the same desired state when its inputs have not changed. It is not enough for terraform apply to exit successfully. The team should be able to run a plan after an apply and explain why it is empty, or why a remaining change is expected.
Use a small convergence loop in CI and during a pre-production rehearsal:
- Validate and format the configuration.
- Create a plan from the intended variable set and provider lock file.
- Review the plan for replacements, access changes, and destructive operations.
- Apply through the approved identity.
- Run a second plan with the same inputs.
- Compare the result with service-side observations, such as cluster reachability and the existence of required topics.
The second plan is a diagnostic artifact. A non-empty result can mean a real configuration change, an eventually consistent API, an omitted argument, a provider default that is not represented in configuration, or a resource changed outside Terraform. Record which explanation applies. Do not suppress a diff with ignore_changes until the team can describe what drift is being accepted and how it will be detected elsewhere.
Lock the provider version for each run and test upgrades against a disposable or staging state before changing production. A provider can expose an implicit default or alter a diff; review the resulting plan as evidence.
5Treat drift as an incident signal, not a nuisance
Drift has several causes, and each calls for a different response. A human may have changed a cluster in the console during an incident. A provider may normalize an API value after creation. An external controller may own a field that Terraform also reads. A failed apply may have created one dependency but not the next.
Classify a diff before correcting it:
| Diff class | Typical evidence | Safe next action |
|---|---|---|
| Intended change | A reviewed commit and matching variable update | Apply from the reviewed branch |
| Out-of-band change | Audit log entry without a Terraform change | Revert through code or record an approved exception |
| Provider normalization | Same service value rendered differently after refresh | Check provider behavior; avoid blind replacement |
| Partial apply | State contains one layer but a dependent layer is absent | Re-plan from the surviving state and complete dependencies |
| Unknown ownership | Another controller changes the field repeatedly | Assign one owner or split the resource boundary |
The evidence is the explanation attached to the diff. Keep the actor, timestamp, resource, old value, intended value, and decision so an on-call engineer can distinguish an unsafe rollback from a representation change.
6Rollback needs a gate before production apply
Terraform can reverse a configuration change, but it cannot restore a Kafka workload's business state by itself. A rollback plan should therefore separate infrastructure reversal from data recovery. Reverting a cluster shape may not recover records removed by a retention policy, and recreating a connector may not undo records already written downstream.
Before a production apply, confirm these gates:
- The previous provider lock file and module inputs are stored with the release.
- A plan shows whether the rollback is an in-place update or a replacement.
- Client endpoints, identities, and topic contracts remain reachable after the change.
- A restore or replay procedure has an owner and a tested starting point.
- The team knows which state snapshot and cloud audit records it will preserve.
Use a staged rollback when the plan includes replacement, access-policy changes, or a network boundary. Apply the smallest reversible change first and keep the original configuration available until the workload's recovery window closes.
7Where AutoMQ fits
The same lifecycle contract applies when the Kafka-compatible data plane uses a different storage architecture. AutoMQ is a cloud-native streaming platform that keeps the Kafka protocol and uses a Shared Storage architecture, with durable stream data stored in object storage and broker compute separated from that storage layer. Its architecture overview and Google Cloud GKE deployment guide identify the environment and path that a team should verify for its chosen product form.
That description does not make an AutoMQ resource interchangeable with a Google-managed Kafka resource. The Terraform provider, module, or deployment API must be checked for the specific AutoMQ BYOC or AutoMQ Software path under review. Keep the same evidence: resource ownership, IAM boundary, network path, secret handling, second-plan convergence, and rollback prerequisites. If the deployment is customer-owned, the cloud account and data-plane boundary should be visible in the graph rather than hidden behind a generic module.
The useful comparison is therefore architectural. A team can keep Kafka clients and partition semantics while choosing how broker compute, WAL storage, object storage, network, and control-plane resources are managed. That choice still needs a provider version, a state boundary, and a runbook that can explain drift.
8FAQ
8.1Does Terraform make GCP Kafka changes idempotent automatically?
Terraform converges on the state its provider can observe and manage. Idempotence still depends on stable inputs, a locked provider, clear ownership, and a second-plan check. Console edits and external controllers can create drift even when the configuration file is unchanged.
8.2Should topics and ACLs share the cluster state?
They can when the same team owns their lifecycle and the provider supports the required resources. Split them when application teams need independent releases or when topic changes require a different approval path. The boundary should follow ownership and recovery responsibility.
8.3Can a rollback restore Kafka records?
No. Terraform can restore infrastructure configuration when the provider supports that operation. Record recovery requires a retention, backup, or replay procedure that the workload team has tested separately.
8.4How should I evaluate managed Kafka Terraform examples?
Use examples to learn current resource syntax, then verify the provider version, IAM permissions, network path, secret handling, and replacement behavior in your own project. A successful sample apply is not evidence that the second plan is empty or that rollback is safe.
The first GCP Kafka Terraform review should end with a graph and a second plan that a different engineer can explain. If identity, network, drift, and rollback still live only in tribal knowledge, make those dependencies part of the contract before adding more resources. If you are evaluating a Kafka-compatible shared-storage deployment on Google Cloud, start an AutoMQ evaluation with the same graph, state policy, and rollback gates.
