QPS & Throughput Playbook
The formula
QPS = (DAU × actions per user per day) ÷ 86,400 → peak QPS = QPS × peak factor (2–10×)
Read vs write
Split traffic into reads and writes — they hit different systems and the ratio drives your whole design. A read:write ratio of 100:1 (typical of social feeds) means caching and read replicas dominate; a write-heavy 1:1 (logging/metrics) means partitioning and write throughput dominate.
Steps
- Average QPS from DAU × actions ÷ 86,400.
- Apply a peak factor (daily peaks ~2–3×; spiky events like flash sales 10×+).
- Split by read:write; size each path separately.
Worked example
50M DAU, 20 actions/day, 90% reads:
- Total = 50M × 20 ÷ 10⁵ = ~10,000 QPS avg (the 10⁵ shortcut — true 86,400-s math gives ~11.6K, the ~14% rounding trap below; fine here because peak ×3 dominates the error) → peak ×3 ≈ 30–35K QPS.
- Reads ≈ 27K QPS, writes ≈ 3K QPS at peak.
- Gate: 27K read QPS ⇒ cache + read replicas; 3K write QPS ⇒ a single primary is still fine (see Decision Gates).
When the QPS formula misleads
| Trap | Wrong conclusion | Fix |
|---|---|---|
| DAU × actions/day / 86400 | Average QPS looks tiny; you size for average and die at peak | Use peak hour share (e.g. 20% of daily in 2h) or measured p99 peak |
| One number for all endpoints | Feed read and like-write share one pool | Split read QPS vs write QPS; different stores and caches |
| Ignoring fan-out | 1 user post → 1 write | Push fan-out can be 1 write × N followers of internal QPS |
| Equating QPS with connections | "10k QPS needs 10k servers" | Little's Law: concurrency = QPS × latency; size threads/conns from that |
Little’s Law, worked: 10K QPS hitting a DB at p50 20 ms holds N = 10,000 × 0.020 = 200 connections on average — but at the p99 of 200 ms the same QPS pins N = 2,000. A 500-connection pool that comfortably fits the mean saturates the moment latency degrades 2.5×: this is why connection pools, not CPU, are the first thing to melt during a slow-DB incident — and why you size pools off degraded-latency concurrency (QPS × timeout budget), not off healthy p50. For why latency degrades nonlinearly as utilization climbs, see the M/M/1 section of the CPU & Server-Count Playbook.
Boundary A (diurnal peak): 10M DAU, 20 posts/day average → 10e6×20/86400 ≈ 2.3k avg write QPS. If 25% of posts land in 2 evening hours: peak ≈ 10e6×20×0.25/(2×3600) ≈ 6.9k write QPS — 3× average. Size for peak.
Boundary B (fan-out + hot shard): 10k frontend QPS with fan-out 50 → 500k backend QPS. Spread over 20 shards ≈ 25k each on average — but if one celebrity key draws 2% of reads, that shard carries ~ (0.98×500k)/20 + 0.02×500k ≈ 34.5k QPS while peers sit near 24.5k. Average capacity passes; one shard melts. Always apply fan-out before per-node gates, then check concentration.
Rounding trap: 600M req/day ÷ 100,000 ≈ 6,000 QPS, but true avg is 600M/86,400 ≈ 6,944 (~14% undercount) before any peak factor.
See also: Senior corrections.
Interviewer follow-ups
- Why not just auto-scale on CPU? CPU can be fine while queue lag or DB connections are the bottleneck — pick the SLI that matches the work.
- Is 3k write QPS fine on one primary? Only after fan-out and hot-key checks; 3k frontend writes can become 3k×N internal writes.
Formulas are standard/public-domain engineering math. Approach and reference-table format adapted from the System Design Primer (CC BY 4.0), Jeff Dean’s latency numbers, the DesignGurus capacity-estimation guide, and Little’s Law.
🤖 Don't fully get this? Learn it with Claude
Stuck on QPS & Throughput Playbook? Open Claude, copy a block below, and it'll teach you this exact concept — visually and interactively.
Build the mental picture, not memorization.
I just read a lesson on **QPS & Throughput Playbook** (System Design) and want to truly understand it. Explain QPS & Throughput Playbook from first principles using ONE vivid real-world analogy and a visual mental model — draw it as ASCII art or a clear step-by-step diagram — with a concrete example using real numbers. Then ask me one question to check I got the mental picture, and wait for my reply. If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.
Socratic — adapts to where you're stuck.
Teach me **QPS & Throughput Playbook** interactively. Ask me ONE guiding question at a time, wait for my answer, and adapt to my confusion — build the idea with me step by step instead of explaining it all at once. If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.
Active recall exposes what you missed.
Quiz me on **QPS & Throughput Playbook** with 5 questions, easy to tricky, ONE at a time. Tell me if each answer is right; at the end, explain clearly what I got wrong and why. If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.
Intuition + hook + flashcards for long-term memory.
Help me remember **QPS & Throughput Playbook** for the long term: give the one-sentence intuition, a memorable hook/mnemonic, a tiny worked example, and 3 active-recall flashcards (Q -> A). If you're unsure or a claim isn't standard, say so and reason from first principles instead of guessing.