System Design

Back-of-the-Envelope Estimation for System Design

How to do back-of-the-envelope estimates for system design: useful numbers, QPS, storage and bandwidth maths, latency orders of magnitude and a worked example.

A worn paper envelope with pencil sketches of boxes and arrows on a dark desk beside a calculator
Illustration: Backend Architect / AI-generated.

Key takeaways

  • Estimates exist to guide design decisions, so round aggressively and state your assumptions.
  • Remember a few anchors: about 100,000 seconds in a day, and a million requests a day is roughly 12 per second.
  • Estimate traffic, storage, bandwidth and memory, then sanity-check against what one machine can do.
On this page

Back-of-the-envelope estimation means working out rough numbers, such as requests per second, storage per year or bandwidth at peak, quickly enough to guide a design. The aim is the right order of magnitude, not precision: knowing whether you need one database or fifty, and whether data fits in memory, is what matters. Round hard, write your assumptions down and check the result against common sense.

Why it matters

Estimates decide architecture. A service handling 50 requests per second runs happily on a single server with a replica; one handling 500,000 needs caching, partitioning and careful capacity planning. In interviews, estimation shows you can connect requirements to design; in real projects, it prevents both over-engineering and nasty surprises. Our system design interview framework places it right after requirements.

Numbers worth remembering

QuantityHandy approximation
Seconds in a day86,400, call it 100,000 (10⁵)
Seconds in a monthAbout 2.5 million
1 million requests a dayAbout 12 per second
1 billion requests a dayAbout 12,000 per second
Peak trafficOften 2–3× the daily average, more for spiky products
1 KB × 1 million1 GB
1 MB × 1 million1 TB

Powers of two help too: 2¹⁰ ≈ a thousand (KB), 2²⁰ ≈ a million (MB), 2³⁰ ≈ a billion (GB), 2⁴⁰ ≈ a trillion (TB).

Latency orders of magnitude

Peter Norvig’s widely quoted table of approximate timings makes the key point: each layer is orders of magnitude slower than the one before.

Operation (Norvig’s approximate figures)Time
Fetch from main memory~100 nanoseconds
Send 2 KB over a 1 Gbps network~20 microseconds
Read 1 MB sequentially from memory~250 microseconds
Disk seek (spinning disk)~8 milliseconds
Read 1 MB sequentially from disk~20 milliseconds
Packet from the US to Europe and back~150 milliseconds

Those figures come from older hardware: modern memory is faster, and SSDs are far faster than spinning disks for random reads. What stays true: memory is much faster than disk, local calls are much faster than network calls, and cross-region round trips are expensive. That’s why caching and keeping chatty calls inside one region matter so much.

A method that works

  1. State the assumptions: daily active users, actions per user, read-to-write ratio, object sizes, retention period.
  2. Traffic: requests per day ÷ 100,000 ≈ average per second; multiply for peak.
  3. Storage: new objects per day × size × retention, then add replication (often ×3).
  4. Bandwidth: requests per second × response size, separately for reads and writes.
  5. Memory for caching: if 20% of objects serve most reads, how big is that 20%?
  6. Sanity-check: compare with what one machine or one database can handle, and design from there.

Worked example: a URL shortener

Take the service from our URL shortener walkthrough:

  • Assumptions: 100 million new links a month; a 100:1 read-to-write ratio; about 500 bytes per record; keep links for 5 years.
  • Writes: 100 million ÷ 2.5 million seconds ≈ 40 per second on average.
  • Reads: 100 × 40 = 4,000 redirects per second on average; plan for perhaps 10,000 at peak.
  • Storage: 100 million × 500 bytes = 50 GB a month; × 60 months = 3 TB; × 3 replicas ≈ 9 TB.
  • Cache: caching the hottest 20% of a month’s links is about 10 GB, which fits comfortably in memory.
  • Conclusion: writes are modest; reads dominate, so a cache in front of a replicated store handles them, and sharding can wait until the data outgrows one cluster.

Worked example: chat messages

For a chat system with 50 million daily users sending 40 messages each:

  • 2 billion messages a day ≈ 20,000–25,000 per second on average, perhaps 60,000–70,000 at peak.
  • At about 200 bytes each, that’s 400 GB a day, or roughly 150 TB a year before replication.
  • The conclusion follows quickly: messages need a horizontally scalable store, partitioned by conversation.

Common mistakes

  • False precision. 86,400 is right, but 100,000 is easier and just as useful.
  • Forgetting peaks. Systems fail at peak, not on average.
  • Ignoring replication and indexes, which can double or triple storage.
  • Mixing bits and bytes in bandwidth maths (network speeds are usually quoted in bits).
  • Not stating assumptions, which makes the numbers impossible to check or adjust.

Frequently asked questions

How accurate do the estimates need to be?

Within a factor of a few is usually fine. The point is to pick the right design, such as one database versus a sharded cluster, not to predict an exact bill.

Should I memorise latency numbers?

Memorise the orders of magnitude, not the digits: nanoseconds for memory, sub-millisecond round trips inside a data centre, around 100 milliseconds across continents.

What if my assumptions are wrong?

Say so, and show how the design changes if traffic is ten times higher. That flexibility is often the most valuable part of the exercise.

Sources

  1. Peter Norvig — Teach Yourself Programming in Ten Years (approximate timings)
  2. Google — Site Reliability Engineering books

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading