Table of Contents
Table of Contents
A document that proposes an Apache Kafka platform change rarely fails on its architecture. It fails because the two groups that must approve the change, the engineering leaders and the finance reviewers, read the same page and answer from different checklists. Two migration proposals crossed the same desk last quarter, one written by an architect and one by a finance lead. The first led with a broker topology diagram and a list of protocol features. The second led with a savings estimate that never stated its assumptions. Both documents were rejected, and both rejections were described as "a review-panel decision," the polite name for a room that stopped trusting the document.
The room is predictable. A CTO review is not a deeper technical audit; it is a fixed set of questions asked in a stable order: what problem exists, what the change must achieve, which options were compared on the same criteria, what the money means, and what happens if the plan is wrong at month three. A Kafka business case built around those five questions survives because the room never has to reconstruct the author's intent. The format that keeps showing up on the winning side of those reviews has five sections, one per question.
1Why good architecture loses the review
The architecture-led document was, technically, correct. It named the right broker count, the right replication design, and the right retention periods. It failed because it answered questions nobody in the room had asked. A CTO is not deciding whether the architecture is sound; that judgment usually happened weeks earlier in an engineering forum. The review decides whether the change is worth the organization's budget, attention, and risk. Those are financial and operational questions wearing a technical coat.
The room asks three overlapping sets of questions. The architecture lead asks what is broken and what was compared. Finance asks what the change costs and when the money returns. The CTO asks the synthesis question that decides the room: if the plan is wrong at month three, what has the team bought and how does it stop spending. The diagram maps each question to the section that carries its answer.
2Five sections that answer the room's questions
The surviving format has five parts in a fixed order, and the order mirrors how trust rebuilds in a review. Problem evidence comes first, then the goal, then options scored on shared criteria, then financials, then risk with phased investment. A case that opens with cost invites a fight about the estimate before the problem is agreed. A case that opens with architecture invites a feature-by-feature tour that walks past the question the room came to settle.
Each section has one job, and that constraint is what stops a business case from becoming a feature tour. Problem evidence names the failure and the metric that proves it. The goal states the end state as a measurable number. Options is a scoring table where every candidate, including the do-nothing baseline, faces the same four criteria: compatibility, cost, operations, and exit. Financials converts each option into a three-year line, and risk stages the spend behind decision gates. The ordering rule matters as much as the sections themselves; when cost appears before evidence, the room reopens the assumptions and debates the arithmetic instead of the problem.
| Candidate | Compatibility | Cost model | Scaling model |
|---|---|---|---|
| Self-managed Apache Kafka | Full protocol control, full operational load | Broker fleet, attached disks, and replica traffic | Add brokers and rebalance partitions |
| Managed Kafka service | Protocol plus a provider console | Provisioned throughput plus retention | Scaling within provider tiers |
| Object-storage-native platform (AutoMQ) | Kafka-compatible | Compute plus Shared Storage | Stateless brokers with separation of compute and storage |
If the shortlist includes an object-storage-native platform, one entry worth scoring is AutoMQ, a Kafka-compatible streaming platform. Score it on the same four criteria rather than on its storage design. Its compatibility with Apache Kafka and architecture overview are the two references to attach to the scoring table before the review, so the compatibility and cost rows trace to a source instead of a claim.
3Translating latency and partitions into money
Finance does not read partitions; it reads line items. The bridge is a translation table that carries each technical signal to a dollar consequence, and the table is the fixed part of the format while the numbers are re-derived from each team's own workload. The five rows that matter are latency, partition count, broker count, replication factor, and rebalances, because those are the engineering signals a CTO will hear about later in an incident review.
Work one example out loud so the room can follow the arithmetic, then keep the do-nothing baseline on the same grids while it happens. Assume a workload writes 100 GB of logical producer traffic per day with a replication factor of three spread across three Availability Zones (AZs). Assume each follower copy crosses one AZ boundary, and the cloud meter charges $0.02 per GB, or $20 per TB, of cross-AZ data transfer. Each logical gigabyte then produces two follower copies, or 200 GB of cross-AZ movement per day.
- Price the month. 200 GB of movement per day is 6 TB per month; at the assumed rate, replication adds $120 a month, or $1,440 a year, for each 100 GB per day written.
- Price the retention. The same workload at three copies and 90 days of retention holds 27 TB of durable data, before compression.
- Keep the baseline honest. The do-nothing option lands on the same grids at the same unit rates, so every delta stays visible instead of floating.
Every number above traces to an assumption stated in the previous paragraph, which is what lets the CTO challenge one line instead of throwing out the whole case. Partition count works the same way and follows the one-step rule. State the memory and CPU each partition consumes on a broker, add the rebalance overhead at the proposed count, and the resulting broker hours become the finance line.
A technical signal that cannot be traced to a dollar line in one step belongs in the engineering appendix, not in the financials section.
4The three fatal omissions
The same three omissions appear in most rejected cases, and each one reads to the room as an incomplete comparison rather than a wrong number. Name them before the review, and the discussion moves from whether the case is honest to whether the assumptions are right. That second discussion is the one a proposer can win with evidence.
- No baseline. Symptom: the proposal compares the replacement platform against a cost of zero. Why it fails: finance has no counterfactual, so every savings figure is unanchored. Fix: state the current platform's three-year run cost with the same line items used for every option.
- No exit cost. Symptom: the migration is priced as if the old platform disappears on cutover day. Why it fails: the old platform stays on the books through the retention and rollback window, and that overlap is usually the largest forgotten line. Fix: add a ramp-down line that covers dual-run, rollback capacity, and decommissioning labor.
- One-shot investment. Symptom: the case asks for the full amount before any evidence exists. Why it fails: the room is asked to bet the whole budget on assumptions a pilot has not tested. Fix: split spend into phases, each ending at a decision gate with an exit condition.
Fixing the omissions rarely changes the architecture in the case. It changes the shape of the ask. A phased case with a named baseline and an exit line gives the CTO something to approve incrementally, and incremental approval is what a multi-year platform change can survive.
5A template you can fill in an afternoon
The fillable form below is the working version of the five sections. Each block has a prompt and a stopping rule. Nothing in it requires a vendor meeting or a benchmark run. Fill it with numbers the team already owns, and mark every cell that still needs a source.
- Problem evidence. Name the failure in one sentence, attach the metric that proves it, and state the affected workflow and its owner.
- Goal. Write the end state as a measurable outcome, for example "hold read tail latency within budget at three times current volume with the same headcount."
- Options. List every candidate, include the do-nothing baseline as a first-class row, and score each on the same four criteria.
- Financials. Give every option a three-year line covering compute, storage, network, people, and exit. Show the arithmetic for at least one line the way the replication example does.
- Risk and phased investment. Split spend into phases, each with a decision gate, and name the failure each phase is buying evidence about.
The single rule that makes the format work is common criteria: hold the columns fixed even when a vendor's material argues for its own rows. If each option is scored on its own terms, the table becomes a set of anecdotes. Fix the columns first, and the scoring stops being a negotiation about what the words mean.
The two proposers from the opening each brought a second draft to last month's review, weeks before the room reconvened. The architecture-led draft became a problem-evidence section, and the savings-led draft became a financials section with an assumption line. Neither team changed its design, and both cases passed, because the room could finally see which facts supported the ask. Run the same scoring exercise against AutoMQ and start with the translation table; the business case a CTO can audit is the one a CTO approves.
6References
- Apache Kafka documentation
- Apache Kafka KIP-405: Kafka Tiered Storage
- AutoMQ compatibility with Apache Kafka
- AutoMQ architecture overview
- The inter-AZ traffic multiplier in Kafka
- Sizing Kafka partitions to cost
- Measuring cost per streaming gigabyte
7FAQ
7.1What is the minimum a streaming platform business case must contain?
A defensible case needs the five sections in order: problem evidence with the metric that proves the failure, a measurable goal, options scored on the same criteria, financials with a stated assumption line, and risk with phased investment. Everything else, feature tours, vendor scorecards, benchmark charts, belongs in an appendix. The room approves or rejects on the five sections alone.
7.2When does the financials section fail?
It fails when finance cannot trace a single number back to its assumption. A savings figure without a baseline, a network line without a crossing rule, or a missing exit cost all collapse the section, because the CTO then has to audit the author instead of the arithmetic. Every dollar line needs a one-step trace to a stated input.
7.3How do you convert latency into a dollar figure a CTO accepts?
Trace the latency signal through its operator. State the traffic covered by the tail, the revenue or worker time that depends on it, and the fraction of that value a latency budget change puts at risk. The number is smaller and more credible than a headline revenue figure, which is exactly why the room trusts it.
7.4Does the format change for an object-storage-native candidate?
The format does not change; the scoring rows change. An object-storage-native platform is scored on compatibility, cost, operations, and exit like every other row, with its compatibility and architecture documents attached as the source for those scores. The neutrality of the format is the reason a CTO can compare it against a self-managed or managed Kafka option.
7.5How long should the review document be?
Long enough to carry the five sections and nothing more. A three-page case with a traceable assumption line beats a forty-page case with a benchmark appendix, because the review is a decision about money and risk, not a feature audit. Keep the format fixed and cut anything that does not answer one of the room's five questions.
