Blog

Kafka in Your VPC or Someone Else's: A Data Control and Network Trade-Off Review

Table of Contents

Table of Contents

A security review can approve a private Kafka endpoint and still leave the most important question unanswered: where does the record live after the client sends it? The endpoint may sit inside your VPC while the Kafka data plane, durable storage, operational metadata, logs, and support path sit in another account. A private route protects one connection. It does not define the whole service boundary.

For teams evaluating Kafka BYOC vs SaaS, that is the fault line. BYOC places the service environment in the customer's cloud account and network boundary. SaaS Kafka places the data plane in infrastructure operated by the service provider. Neither label settles the security or compliance decision by itself. The useful comparison is more precise: which data classes stay in your account, which paths leave it, who can inspect the system during an incident, and what evidence can your team produce later?

A sound review keeps three ideas separate. Data control describes where records and related resources live and who controls the account, storage, keys, and routes. Service responsibility describes who runs upgrades, scaling, incident response, and support. Compliance evidence describes whether the resulting controls satisfy a specific requirement. Confusing those layers is how a team ends up saying “the traffic is private” when it has verified a single network hop.

SaaS Kafka and BYOC Kafka data paths across cloud account and VPC boundaries

1Data control starts with a complete inventory

A Kafka data-control review should begin with an inventory, not a deployment label. Records are the visible class, but they are not the full set of information that can reveal a workload or affect a residency decision. Topic and partition metadata, consumer offsets, schemas, ACLs, audit events, metrics, logs, backups, and encryption keys may follow different paths.

The inventory should answer four questions for each class of data:

  • Where is it stored? Name the cloud account, region, VPC, bucket, disk, database, or managed service.
  • Who controls it? Identify the account owner, storage policy owner, key administrator, and identities that can change access.
  • How does it move? Record client traffic, replication, object storage access, telemetry, support access, migration traffic, and disaster recovery paths.
  • What proves the answer? Point to a deployment manifest, cloud resource inventory, flow log, access policy, audit record, or tested runbook.

This inventory separates a useful claim from a vague one. “Data stays in our VPC” can mean records stay there while logs leave, or it can mean the data plane, storage, and operational metadata all run inside the same customer-controlled environment. The difference matters to residency, incident response, and contract review.

SaaS can still meet a team's requirements. A provider may offer regional placement, encryption, private connectivity, access logs, and a well-defined support process. The buyer's task is to verify those controls against the workload and retain the evidence. BYOC can make the cloud account boundary more direct, but the customer then owns the cloud configuration that makes the boundary hold.

2Follow the record, then follow the side paths

In a SaaS deployment, a producer in the customer VPC usually reaches the service through a provider-supported endpoint or network connection. The Kafka brokers and durable storage run in provider-managed infrastructure. The client path can be private while the record still crosses from the customer account into the provider's account. Consumer reads, connector traffic, monitoring export, and support actions can take separate routes.

In a BYOC deployment, the data plane runs in the customer's cloud account and VPC. Client traffic, broker coordination, and storage access are designed within that environment, subject to the provider's deployment model and the customer's route, security group, endpoint, and identity policies. That can reduce the number of account boundaries in the main data path. It does not make every path local by default.

Private networking has a narrower meaning than many architecture diagrams suggest. A private endpoint can keep client traffic off the public internet, but it does not erase endpoint processing charges, endpoint hourly charges, NAT traversal, cross-AZ traffic, inter-region transfer, or traffic to a connector in another VPC. The route still has a source, destination, zone, account, and billable cloud service. Review each path on those terms.

For a network review, draw these flows separately rather than using one arrow labeled “Kafka traffic”:

  1. Producer and consumer traffic to Kafka brokers.
  2. Broker access to durable storage and write-ahead log storage.
  3. Schema, connector, stream-processing, and sink traffic.
  4. Metrics, logs, audit events, and backup or replication traffic.
  5. Vendor maintenance, support, and emergency troubleshooting traffic.

The last two flows often determine whether a residency statement is complete. A platform can keep records in one region while exporting logs or support diagnostics elsewhere. That may be acceptable, but it must be governed.

For a deeper look at how endpoint placement changes the design, see Private Networking for Kafka-Compatible Data Planes. The useful question is not whether a connection is called private. It is whether every meaningful path has an owner and a reviewable policy.

3The control surface changes when Kafka moves accounts

A SaaS service gives the provider control over the infrastructure substrate. That can remove work from the platform team, but it also means troubleshooting and audit evidence depend on the provider's visibility, access model, and service interfaces. BYOC changes the control surface. The customer owns the cloud account and its policies, while the provider may still operate the Kafka software through delegated permissions or an approved support channel.

The comparison is easier when the review uses the same questions for both models:

Review questionSaaS KafkaBYOC Kafka
Where does the data plane run?Provider-managed account or service boundaryCustomer cloud account and VPC, according to the deployment model
Who controls cloud network policy?Provider, with customer-controlled client-side pathsCustomer, including VPC, routes, endpoints, and security groups
Who can inspect a production failure?Provider operators and customer operators through the service interfaceCustomer cloud operators plus provider access granted for management or support
Where do audit records come from?Provider service logs plus customer-side identity and network logsCustomer cloud audit logs plus platform and provider access records
Who bears the configuration risk?Provider for the service substrate, customer for clients and integrationsCustomer for cloud foundation and permissions, shared with provider for the service layer

The table identifies where proof must come from. A SaaS review asks for provider evidence and tests the customer-side path. A BYOC review asks for the same service evidence plus the customer's IAM, storage, VPC, endpoint, and logging configuration.

Kafka data control evidence matrix for SaaS and BYOC deployment models

Troubleshooting access deserves its own review. During a broker incident, can the provider see the required metrics and logs without reading payloads? Can the customer inspect flow logs and cloud audit events? Is support access time-limited, approved, and recorded? Does emergency access use a separate identity from routine control-plane access? These questions determine whether “managed” means reduced toil or an unexamined access path.

Data control and service responsibility can move in opposite directions. A customer may keep the data plane in its account while delegating maintenance. Another customer may retain data in a provider service but require strict approval for support actions. The decision is sound when the contract, architecture diagram, IAM policy, and incident runbook tell the same story.

4Private networking has a real cost boundary

Private connectivity is a control mechanism first. It can reduce exposure and fit internal routing policy. It is not a promise that the network bill disappears. The cost depends on the path, the cloud provider, the endpoint type, the availability-zone layout, the volume of bytes processed, and the location of clients and sinks.

A SaaS design can add a private endpoint between the customer VPC and the provider service. The customer may still pay for endpoint resources and data processing, while the service may apply its own network or usage dimensions. A BYOC design can use gateway or interface endpoints for cloud services, but those endpoints, NAT gateways, load balancers, and cross-VPC routes remain customer resources. Compare the actual topology and current provider pricing pages, not the words SaaS or BYOC.

Keep these questions beside the network diagram:

  • Does the client path cross an Availability Zone before reaching the endpoint or broker?
  • Does object storage use a gateway or interface endpoint, and which route tables cover it?
  • Do connectors and stream-processing jobs run in the same VPC and region as Kafka?
  • Does support or telemetry leave the approved account or region?
  • Which cloud account receives each endpoint, NAT, load balancer, and transfer charge?

This is where data control meets FinOps. Keeping Kafka in your VPC can improve account-level visibility, but it can also expose cloud costs that a SaaS invoice previously bundled into a service charge. That changes observability, not necessarily cost.

5BYOC can clarify the boundary without answering compliance for you

The architectural requirement is now clear: a team that needs customer-account placement should be able to keep the data plane, durable storage, metadata services, and network controls inside an approved cloud boundary, while making provider responsibilities explicit. That is the point at which a concrete BYOC implementation belongs in the discussion.

AutoMQ BYOC is one factual example. AutoMQ's BYOC documentation states that the environment's physical resources belong to the user, and that the Control Plane and Data Plane are deployed in the user's network. Its AWS preparation guide states that all BYOC components are deployed within the customer's AWS account. In that model, Kafka records and the metadata managed by the deployed Kafka components are part of the customer-account boundary, while logs, metrics, support access, and cloud service calls still need to be mapped in the review.

That placement is useful when the platform team needs its own VPC, IAM, storage policies, cloud audit trail, and cost allocation. It does not turn AutoMQ BYOC into a self-managed cluster. The provider can still support environment maintenance and version upgrades through customer-authorized access. The shared model is explicit: the customer governs cloud resources and permissions, and the provider operates the service layer within the granted scope.

The architecture also affects the network discussion. AutoMQ uses a Kafka-compatible data plane with shared storage and object-storage access paths. On AWS, the preparation guide calls for an S3 Gateway Endpoint and an EC2 Interface Endpoint for the deployment, along with private subnets, DNS, and route configuration. Those requirements make the network review concrete. They show why BYOC still requires network engineering, while putting the decisions in the customer's account for inspection and governance.

For teams comparing data ownership across regions, Data Ownership Trade-Offs in Multi-Region Private Networking provides a related design lens. Document each data class, its path, and the control that permits it.

6A decision table for Kafka in your VPC

“SaaS fit” and “BYOC fit” describe conditions, not universal outcomes. The evidence column is the artifact to request before a decision is recorded.

RequirementSaaS fitBYOC fitEvidence to request
Records and metadata must stay in a customer cloud accountWeak unless the service offers that placementStrong when the deployment places the relevant components and storage thereAccount map, storage inventory, metadata path, and deployment documentation
Client traffic must use private routesSupported through the provider's private connectivity optionsDesigned through customer VPC, DNS, routes, endpoints, and security groupsSubnet-level data-flow diagram and flow-log test
The platform team needs cloud-level troubleshootingDepends on provider logs, APIs, and support accessCustomer can inspect cloud resources, with provider access defined separatelyIAM policy, audit logs, support runbook, and incident exercise
The organization needs regional data residency evidenceProvider documentation and service controls must be sufficientCustomer controls more of the region and resource placementRegion map, storage policy, telemetry path, and deletion process
The team wants low infrastructure ownershipUsually a stronger SaaS fitPossible, but cloud foundation remains a customer responsibilityResponsibility matrix and operating calendar

Before choosing, ask the service owner to sign off on five artifacts: the data-class inventory, network path map, access and support model, cloud cost boundary, and exit plan. The exit plan should cover records, schemas, offsets, ACLs, connectors, audit evidence, and customer-owned storage that must remain readable.

Kafka deployment decision table for customer-account and provider-account requirements

A compliance review should then name the actual requirement. “The data is in our VPC” is an architecture fact. “The deployment meets our regulatory obligation” is a conclusion that depends on the regulation, data classification, controls, contracts, and evidence. Keep those statements separate so an auditor can test each one.

7FAQ

7.1Does a private Kafka endpoint mean the data stays in my VPC?

No. It describes the client connection path. Confirm where brokers, durable storage, metadata, logs, metrics, support diagnostics, and connectors run. A private endpoint can connect your VPC to a provider account.

7.2Is BYOC Kafka the same as self-managed Kafka?

No. BYOC places the service environment in the customer's cloud account, while the provider may continue to manage software lifecycle tasks under a shared responsibility model. Self-managed Kafka usually leaves the customer responsible for the full platform stack.

7.3Does BYOC guarantee compliance or data residency?

No. BYOC can give the customer more direct control over account, region, storage, network, and IAM boundaries. Compliance and residency conclusions still require a requirement-specific review of data classes, controls, contracts, access, telemetry, and deletion.

7.4What should a Kafka data-control review document?

Document records, topic and partition metadata, offsets, schemas, ACLs, logs, metrics, backups, keys, support access, network paths, account ownership, regions, and retention or deletion behavior. Attach an evidence source to each item.

7.5Can private networking still create Kafka egress or endpoint costs?

Yes. Endpoint processing, endpoint resources, NAT, cross-AZ traffic, inter-region transfer, load balancing, and connector placement can create cloud charges. Trace the route and check the current cloud pricing page for the services it uses.

8References

The first question in the review was where the record lives after the client sends it. The next step is to answer that question with your own account map, path test, and access evidence. If the target architecture requires a customer-owned cloud boundary, start an AutoMQ BYOC evaluation with those artifacts ready.

Newsletter

Subscribe for the latest on cloud-native streaming data infrastructure, product launches, technical insights, and efficiency optimizations from the AutoMQ team.

Join developers worldwide who leverage AutoMQ's Apache 2.0 licensed platform to simplify streaming data infra. No spam, just actionable content.

I'm not a robot
reCAPTCHA

Never submit confidential or sensitive data (API keys, passwords, credit card numbers, or personal identification information) through this form.