Table of Contents
Table of Contents
Most platform teams inherit a binary when they choose how to run Apache Kafka: operate it themselves or buy it as a managed service. Each pole of that binary has a trade everyone can recite. Self-managed Kafka hands the team full control and a permanent operations tail: broker replacement, partition rebalances, disk capacity, rolling upgrades, and the on-call rotation that keeps all of it moving. A managed Kafka service removes most of that tail and replaces it with a different cost: service fees, metered storage and transfer, and a boundary past which the team no longer controls the data path. The binary is comfortable because both answers are legible, which is exactly why it hides a third option that changes the axis of the decision.
That third option is the object-storage-native Kafka platform. It keeps the Kafka protocol but moves durability out of broker-local disks and into shared object storage. The broker becomes a stateless compute process; partitions stop being tied to the disk a broker happens to own; capacity changes stop being storage migrations. This is not a midpoint between the two poles. It is a different operating model with its own cost shape, ops burden, and risk profile, and for a meaningful set of teams it is the correct answer rather than the compromise. The decision path below turns the choice into three questions, three profiles, and a sequence of transitions you can walk before entering the review room.
1Not two options, three
The self-managed versus managed debate is loud because both sides are arguing about the wrong variable. The real variable is where durability and partition state live, because that one fact determines what work the team keeps, what it hands off, and what it pays for. In a self-managed deployment the state lives on broker-attached disks, so the team inherits the work of keeping that state balanced, replicated, and recoverable. In a managed service the provider owns that state on the team's behalf, and charges an operating margin plus the underlying infrastructure for it. In an object-storage-native platform the state lives in an object store the team can see and account for directly, while the broker layer shrinks to stateless compute that can be resized without moving data.
Each location of state produces a distinct failure story, because the question asked at two in the morning changes. Self-managed asks: which broker holds the partition, and how many replicas are still in sync? Managed asks: what is the provider's recovery path, and what is the support boundary? Object-storage-native asks: is the object store durable, and are the broker processes behind the protocol healthy? The questions are very different, and a team should know which question it wants to own before it picks a path.
2Three questions that split the decision
Before comparing paths, pin down the team that will run the platform. Three questions separate the decision, and each one is answerable from evidence the team already carries in its backlog and cloud bill.
How many engineer-hours a week can the platform absorb? The answer is not a headcount estimate; it is the number of hours that current tickets already consume. If a team has a dedicated streaming group that runs rebalances and broker upgrades as routine work, the operational tail of self-managed Kafka is a known quantity. If the same team is three people who also own the data warehouse and the CI pipeline, that tail is a recurring incident. Ops capacity below a threshold pushes the decision away from self-managed, regardless of how much the team likes owning the data path.
What is the data shape? Retention depth, partition count, traffic burstiness, and the size of the hot read set matter more than total throughput when choosing a storage model. A workload that keeps months of history with long-tail reads suits object storage, where the bill follows what is stored and fetched rather than what is provisioned. A workload with a small, hot, latency-sensitive set and short retention suits local-disk designs, where the hot data sits close to compute. Painting every workload with the same brush is how teams end up provisioning for worst-case retention they never reach.
Where does it run, and who owns the account? An on-premises or air-gapped requirement collapses the question to self-managed or a software deployment in the team's own data center. A cloud deployment opens the second question: whether the team wants the control plane and data plane inside its own cloud account, or is willing to operate on someone else's metered infrastructure. The first answer points to bring-your-own-cloud models; the second points to the classic managed service.
The same evidence, organized as a table, keeps the split points straight.
| Question | What it measures | Where it points |
|---|---|---|
| Weekly ops capacity | Hours the platform can absorb before incidents | Below threshold: managed or object-storage-native |
| Data shape | Retention, burstiness, hot-set size | Long retention, spiky traffic: object storage; small hot set: local disk |
| Cloud posture | On-premises, own-cloud-account, or outsourced | On-premises: self-managed or software; own account: BYOC; outsourced: managed service |
None of the three questions decides alone. They act together as a filter: a team with weak ops capacity and spiky, retention-heavy traffic should not be choosing between the two poles of the old binary because the object-storage-native profile matches both constraints.
3The profile and price of each path
Each operating model has a distinct profile once the three questions are answered. The profiles matter because the cost of a Kafka deployment is mostly a function of its architecture, not its brand.
Self-managed Kafka is the control play. The team selects broker hardware, tunes every broker and topic setting, and owns the data path end to end. The price of that control is continuous: block storage is provisioned per GiB-month whether or not the bytes are used, brokers are sized for peak rather than average, and every partition rebalance or broker replacement is scheduled work performed by the team's own engineers. Self-managed earns its place when the organization has a real streaming team, an on-premises or air-gapped requirement, or compliance rules that place the full stack under direct control.
Managed Kafka is the time play. The provider operates brokers, replication, and upgrades, and the team keeps clients, topics, quotas, and governance. The price is metered rather than fixed: service fees, storage, and typically per-GiB charges for the cross-AZ replication traffic that a three-replica deployment generates continuously. Teams trade a known ops burden for a pricing structure they must forecast, and for a data path whose internal failure modes live behind a support ticket.
Object-storage-native Kafka is the storage play. Durability moves to object storage, so brokers become stateless compute and a partition is no longer bound to one node's disk. Scaling the cluster means adding or removing broker processes, not moving partition data between disks. Capacity is elastic in both directions because there is no pre-provisioned broker storage to carry. The cost profile follows: storage bills like object storage, on what is stored and fetched rather than what is reserved, and the brokering layer stays small because it no longer warehouses data.
AutoMQ is one Kafka-compatible implementation of the object-storage-native profile. It is fully compatible with Apache Kafka at the protocol level, and its Shared Storage architecture separates compute from storage so the durability layer sits in S3-compatible object storage behind a WAL (Write-Ahead Log) that handles the hot write path. It is one way to reach the profile this section describes, not the sole way, and the decision framework stands whether or not the team chooses it.
4Sensible transitions between paths
Teams very often arrive at the decision already inside one of the profiles, and the practical question is how to move. Some transitions are routine; one or two deserve explicit warning.
From self-managed to managed, move the least critical topics first. Start with dev and staging workloads, keep production on the self-managed cluster while the team learns the provider's quotas, networking, and failover behavior, then promote topic by topic behind recorded consumer offsets. The trap is not the data copy; it is re-creating ACLs, quotas, connectors, and monitoring with the same labels the on-call team already knows.
From managed to self-managed, the driver is usually account control or cost, and the team should treat it as re-acquiring an operations function instead of changing a URL. The provider's guardrails do not follow the data, which means the team inherits upgrades, broker replacement, and capacity planning on day one of the cutover.
From either path to an object-storage-native platform, the move is the closest thing to a lateral transition the ecosystem offers. Because the Kafka protocol is preserved end to end, producers, consumers, and tooling keep their existing client configuration, and the storage redesign happens underneath the protocol layer. The work concentrates in what the platform actually changes: data moves into object storage, partitions decouple from broker disks, and the team re-plans capacity around stateless brokers instead of provisioned disks. The same inventory discipline applies as in any migration: capture consumer group offsets, mirror topics, and hold the old cluster as rollback until the retention window closes.
A transition is also a second chance to ask the three questions again. A team that chose self-managed when it was twenty people may no longer have the streaming specialists it had then, and a team locked into a managed contract may have acquired the account-control requirements that object-storage-native BYOC models exist to serve.
5A decision tree to take to the room
The tree compresses the sections above into a walk someone can run on a whiteboard in ten minutes.
Start with cloud posture, because it gates everything else. If the workload must run on-premises, the choice is self-managed Apache Kafka or a software deployment the team runs in its own data center. If the workload runs in the cloud, ask whether the control plane and data plane must sit in the team's own account, which separates a BYOC deployment from an outsourced managed service.
Then ask the ops question. A team with dedicated streaming engineers can absorb the operational tail of self-managed or the integration work of any platform; a team without that capacity should rule self-managed out entirely, whatever its feelings about control.
Finally ask the data question. Long retention, spiky traffic, and read-heavy history favor shared object storage; a small hot set with strict tail latency leans toward local-disk designs. The object-storage-native profile earns its place when the cloud posture is flexible enough to use it and the data shape rewards it.
Walk the tree for your own workload before you accept either pole of the old binary. The two-way debate is a habit left over from a time when Kafka brokers had to own their disks; the third path is the reason to redo the arithmetic. Start the object-storage-native evaluation with the AutoMQ Open Source project, and bring the three questions into the room with you; the team that can name where its partition state lives is the one that gets to choose the model.
6References
- Apache Kafka documentation
- Apache Kafka replication configuration
- Apache Kafka broker configuration, including min.insync.replicas
- Apache Kafka KIP-405: Kafka Tiered Storage
- AutoMQ compatibility with Apache Kafka
- AutoMQ architecture overview
- Cloud native Kafka architecture for managed Kafka
- Why object-storage-native retention exposes Kafka storage assumptions
7FAQ
7.1Is object-storage-native Kafka the same as tiered storage?
No. Tiered storage keeps recent data on broker disks and offloads older segments to object storage, which means the broker still owns hot state and the team still manages a disk-backed layer. An object-storage-native platform puts durability in object storage from the first byte and keeps a WAL (Write-Ahead Log) on the hot path, so brokers are stateless and do not carry a local copy of the log as their durability source.
7.2Which path costs the least?
The lowest-cost path depends on workload shape, not on a universal ranking. Self-managed costs are dominated by people and provisioned block storage; managed costs follow service fees plus metered storage and transfer; object-storage-native costs follow object storage usage plus a small broker layer. A retention-heavy, spiky workload typically favors object storage, while a small hot workload on-premises can cost the least on hardware the team already owns.
7.3Do we lose Kafka compatibility on the object-storage-native path?
Not in the implementations that preserve the Kafka protocol. An object-storage-native platform such as AutoMQ is fully compatible with Apache Kafka at the protocol level, so existing producers, consumers, and tooling continue to work with their current configuration. The change is in where durability lives, which sits below the protocol.
7.4When does it make sense to move from managed to object-storage-native?
When two conditions appear together: the team needs the control plane and data plane inside its own cloud account, and the workload is retention-heavy or bursty enough that metered managed pricing and provisioned storage stop matching the traffic shape. A Kafka-protocol-compatible platform lets the team make that move without rewriting its client estate.
7.5What is the biggest risk of choosing the wrong model?
Choosing between two poles and missing the third. The cost of a wrong choice is rarely the migration itself; it is living for years with a storage model that was sized for a different workload, while the team continues to apply the binary that produced it.
