The System Design Interview: A Step-by-Step Framework
A step-by-step framework for system design interviews: clarifying requirements, estimates, APIs, data models, high-level design, deep dives and trade-offs.
How to do back-of-the-envelope estimates for system design: useful numbers, QPS, storage and bandwidth maths, latency orders of magnitude and a worked example.

Back-of-the-envelope estimation means working out rough numbers, such as requests per second, storage per year or bandwidth at peak, quickly enough to guide a design. The aim is the right order of magnitude, not precision: knowing whether you need one database or fifty, and whether data fits in memory, is what matters. Round hard, write your assumptions down and check the result against common sense.
Estimates decide architecture. A service handling 50 requests per second runs happily on a single server with a replica; one handling 500,000 needs caching, partitioning and careful capacity planning. In interviews, estimation shows you can connect requirements to design; in real projects, it prevents both over-engineering and nasty surprises. Our system design interview framework places it right after requirements.
| Quantity | Handy approximation |
|---|---|
| Seconds in a day | 86,400, call it 100,000 (10⁵) |
| Seconds in a month | About 2.5 million |
| 1 million requests a day | About 12 per second |
| 1 billion requests a day | About 12,000 per second |
| Peak traffic | Often 2–3× the daily average, more for spiky products |
| 1 KB × 1 million | 1 GB |
| 1 MB × 1 million | 1 TB |
Powers of two help too: 2¹⁰ ≈ a thousand (KB), 2²⁰ ≈ a million (MB), 2³⁰ ≈ a billion (GB), 2⁴⁰ ≈ a trillion (TB).
Peter Norvig’s widely quoted table of approximate timings makes the key point: each layer is orders of magnitude slower than the one before.
| Operation (Norvig’s approximate figures) | Time |
|---|---|
| Fetch from main memory | ~100 nanoseconds |
| Send 2 KB over a 1 Gbps network | ~20 microseconds |
| Read 1 MB sequentially from memory | ~250 microseconds |
| Disk seek (spinning disk) | ~8 milliseconds |
| Read 1 MB sequentially from disk | ~20 milliseconds |
| Packet from the US to Europe and back | ~150 milliseconds |
Those figures come from older hardware: modern memory is faster, and SSDs are far faster than spinning disks for random reads. What stays true: memory is much faster than disk, local calls are much faster than network calls, and cross-region round trips are expensive. That’s why caching and keeping chatty calls inside one region matter so much.
Take the service from our URL shortener walkthrough:
For a chat system with 50 million daily users sending 40 messages each:
Within a factor of a few is usually fine. The point is to pick the right design, such as one database versus a sharded cluster, not to predict an exact bill.
Memorise the orders of magnitude, not the digits: nanoseconds for memory, sub-millisecond round trips inside a data centre, around 100 milliseconds across continents.
Say so, and show how the design changes if traffic is ten times higher. That flexibility is often the most valuable part of the exercise.
Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.
A step-by-step framework for system design interviews: clarifying requirements, estimates, APIs, data models, high-level design, deep dives and trade-offs.
A URL shortener system design: requirements, capacity estimates, API, short-code generation, database choice, caching, analytics and scaling trade-offs.
Microservices or a monolith? The real costs and benefits of each, the modular monolith middle ground, and a practical framework for deciding and migrating.