Cost ($) Estimation Playbook
Why this sets you apart
Almost no guide teaches cost estimation — yet “what does this cost to run?” is a senior/staff-level signal. Turning a design into a rough monthly bill shows you reason about trade-offs, not just feasibility.
The formula
Monthly $ ≈ storage_GB × $/GB-mo + compute_hours × $/hr + egress_GB × $/GB + request_millions × $/M
Rough cloud price anchors (order of magnitude)
| Resource | ~Price |
|---|---|
| Object storage (S3-class) | ~$0.02 / GB-month |
| Block storage (SSD/EBS) | ~$0.10 / GB-month |
| Compute (general VM) | ~$0.05 / vCPU-hour (spot ~3–5× cheaper than on-demand) |
| Egress (data out) | ~$0.09 / GB |
| Managed requests/serverless | ~$0.20–$1 / million |
Worked example — the photo service
~27 PB stored (object), modest compute, 1 PB/month egress:
- Storage: 27 × 10⁶ GB × $0.02 ≈ ~$540K/month.
- Egress: 10⁶ GB × $0.09 ≈ ~$90K/month.
- Compute: serving + thumbnailing 10M photos/day is I/O-heavy — say 200 vCPU steady: 200 × $0.05 × 730 ≈ ~$7.3K/month (spot for the async thumbnail fleet cuts it to ~$2–3K).
- Requests: 10M uploads + ~1B views/day ≈ 31B req/mo; at CDN/request pricing ~$0.20–1/M that is ~$6–31K/month.
- Total ≈ $540K storage + $90K egress + $7K compute + ~$15K requests ≈ $650K/month — and the shape of the bill IS the insight: storage is ~83% of it, so the tiering lever (move cold blobs to archive at ~$0.004/GB) dwarfs any compute optimization.
- Sanity: 2× traffic ≈ +$100K (egress + requests scale) but storage keeps compounding regardless of traffic — this bill grows even if the product stalls.
Why 27 PB and not the 81 PB the replication page insists on?
Because who pays for replication depends on who runs the storage. Managed object storage (S3-class) prices logical bytes — the ×3+ durability copies are the provider’s problem, already inside the $0.02/GB-mo. Self-hosted storage (your own disks/Ceph/HDFS) pays for physical bytes: the same 27 PB costs 81 PB of disks, and the replication page’s warning applies with full force. Rule: price managed storage at the RAW number; price self-hosted at RAW × RF. Getting this wrong mis-bills by exactly RF× — here, $540K vs $1.62M/month, the difference between S3 looking expensive and S3 looking like a bargain.
Normalize it: $650K/mo over, say, 50M MAU is ~1.3¢/user/month — comfortably below any ad ARPU; the same bill over 2M users would be 32¢/user and a business problem, not an engineering one. Per-unit cost is how you answer “is that expensive?” — a raw $650K has no meaning until divided by users or requests.
When cost math misleads
| Trap | Wrong conclusion | Fix / boundary numbers |
|---|---|---|
| Only compute | Misses egress, storage, multi-AZ, managed premiums | Cost stack: compute + storage + network + managed + people/oncall. SaaS example: compute $1.5K vs egress $45K — wrong dominant line. |
| List price forever | Overstate 2–3× | 40 vCPU × $0.05 × 730 ≈ $1,460 on-demand; spot often ~$300–500 for stateless; 1-yr reserved ~40% off. Wrong commitment flips the line. |
| Ignoring idle / non-prod | Dev/stage doubles bill | Include non-prod and multi-region DR. |
| Single-region bill for multi-region design | Understate 2–3× | Active-active 3 regions ≈ 3× compute + cross-region transfer. A $50K single-region design can be $150K+ multi-region before growth. |
| Request-priced services at high QPS | "Serverless is cheap" | 2B req × $0.50–1/M = $1–2K; at 20B/mo the same line is $10–20K — switch model when request line > compute line. |
Sanity: if 2× traffic does not ~2× the bill, you misidentified the dominant cost.
Interviewer follow-ups: "Cheapest design?" — cheapest that still meets latency/error SLO. A design that saves 20% compute and misses SLO is not cheaper after incident cost.
Formulas are standard/public-domain engineering math. Approach and reference-table format adapted from the System Design Primer (CC BY 4.0), Jeff Dean’s latency numbers, the DesignGurus capacity-estimation guide, and Little’s Law.
🤖 Don't fully get this? Learn it with Claude
Stuck on Cost ($) Estimation Playbook? Open Claude, copy a block below, and it'll teach you this exact concept — visually and interactively.
Build the mental picture, not memorization.
I just read a lesson on **Cost ($) Estimation Playbook** (System Design) and want to truly understand it. Explain Cost ($) Estimation Playbook from first principles using ONE vivid real-world analogy and a visual mental model — draw it as ASCII art or a clear step-by-step diagram — with a concrete example using real numbers. Then ask me one question to check I got the mental picture, and wait for my reply. If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.
Socratic — adapts to where you're stuck.
Teach me **Cost ($) Estimation Playbook** interactively. Ask me ONE guiding question at a time, wait for my answer, and adapt to my confusion — build the idea with me step by step instead of explaining it all at once. If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.
Active recall exposes what you missed.
Quiz me on **Cost ($) Estimation Playbook** with 5 questions, easy to tricky, ONE at a time. Tell me if each answer is right; at the end, explain clearly what I got wrong and why. If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.
Intuition + hook + flashcards for long-term memory.
Help me remember **Cost ($) Estimation Playbook** for the long term: give the one-sentence intuition, a memorable hook/mnemonic, a tiny worked example, and 3 active-recall flashcards (Q -> A). If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.