System Design

Design a URL Shortener: A System Design Walkthrough

A URL shortener system design: requirements, capacity estimates, API, short-code generation, database choice, caching, analytics and scaling trade-offs.

A whiteboard sketch of a URL shortener architecture with cache and database boxes
Illustration: Backend Architect / AI-generated.

Key takeaways

  • It’s a read-heavy system: optimise redirects with caching and a fast key-value lookup.
  • Generate short codes with a counter plus base62 encoding, or random IDs with collision checks.
  • Record analytics asynchronously so redirects stay fast.
On this page

The URL shortener is the classic warm-up system design problem because it’s small enough to finish in an interview but touches the fundamentals: ID generation, storage, caching and scale. We’ll follow our system design interview framework.

1. Requirements

Functional: create a short link for a long URL; redirect short links to the original; optional custom aliases and expiry.

Non-functional: redirects must be fast (low tens of milliseconds) and highly available; links must never collide or change destination.

2. Estimates

Assume 100 million new links a month:

  • Writes: 100M ÷ (30 × 86,400) ≈ 40 per second on average
  • Reads at 100:1 ≈ 4,000 per second, with peaks several times higher
  • Storage: ~500 bytes per record × 1.2 billion links a year ≈ 600 GB per year

Conclusion: a read-heavy system where caching dominates.

3. API

POST /api/links   { "url": "https://example.com/very/long", "alias": null, "expiresAt": null }
  -> 201 { "code": "a8Xk2Q", "shortUrl": "https://sho.rt/a8Xk2Q" }
GET  /{code}      -> 301 or 302 redirect

Use 301 if links never change and you want browsers to cache the redirect; use 302 if you need every click to reach your servers for analytics.

4. Generating short codes

Option A: counter + base62. A distributed ID generator issues unique integers; encode them in base62 (a–z, A–Z, 0–9). Seven characters give 62^7 ≈ 3.5 trillion codes. Codes are short and never collide, but they are sequential and guessable.

Option B: random codes. Generate random 7-character strings and check for collisions on insert. Harder to guess; needs a uniqueness check.

Option C: hash the URL. Take a hash and truncate it. Deterministic, but collisions must still be handled.

Most designs choose A with a small random component, or B.

5. Data model and storage

A single table: code (primary key), long URL, created time, expiry, owner. Access is a pure key lookup, which suits a key-value or wide-column store; a relational database with an index works well at moderate scale. See SQL vs NoSQL for the trade-offs, and database sharding for when a single node isn’t enough.

6. High-level design

Clients → load balancer → stateless app servers → cache → database. Popular links sit in a cache using the cache-aside pattern. A small number of links receive most traffic, so a modest cache absorbs the majority of reads.

7. Analytics without slowing redirects

Don’t write to the database on every click. Publish click events to a queue and aggregate them asynchronously. Our comparison of Kafka, RabbitMQ and SQS covers the options.

8. Protecting the service

Apply rate limiting to link creation, scan submitted URLs against malware lists, and expire abandoned links.

Frequently asked questions

How long should short codes be?

Six or seven base62 characters is typical, giving billions to trillions of combinations.

Should the redirect be 301 or 302?

301 is cacheable and faster for users; 302 keeps every click visible to your servers. Choose based on whether you need analytics.

How do you prevent collisions?

Use a unique counter-based ID, or check for existing codes on insert and retry.

Sources

  1. Martin Kleppmann — Designing Data-Intensive Applications

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading