Backend Patterns

Caching Strategies Explained: Cache-Aside, Write-Through and Write-Back

Caching strategies and their trade-offs: cache-aside, read-through, write-through, write-back and write-around, plus eviction policies, TTLs and invalidation.

A dark data centre aisle with rows of server racks and glowing cyan status lights
Illustration: Backend Architect / AI-generated.

Key takeaways

  • Cache-aside is the most common pattern: the app checks the cache, then loads from the database on a miss.
  • Write-through keeps the cache consistent at the cost of write latency; write-back is fast but risks data loss.
  • Most caching bugs come from invalidation; use TTLs, clear keys on writes and protect against stampedes.
On this page

A cache stores copies of data in a faster place, usually memory, so repeated requests don’t hit slower systems like databases or external APIs. Done well, caching cuts latency and load dramatically. Done badly, it serves stale data and creates subtle bugs. The strategy you choose decides which trade-offs you accept.

The read strategies

Cache-aside (lazy loading)

The application manages the cache directly:

  1. Look up the key in the cache.
  2. On a hit, return it.
  3. On a miss, read from the database, store the result in the cache and return it.
def get_user(user_id):
    key = f"user:{user_id}"
    cached = cache.get(key)
    if cached is not None:
        return cached
    user = db.query_user(user_id)
    cache.set(key, user, ttl=300)
    return user

Pros: simple, only caches data that is actually requested, and the app keeps working if the cache fails. Cons: the first request for each key is slow, and cached data can go stale until it expires or is invalidated.

Read-through

The cache sits in front of the database and loads missing data itself. Application code only talks to the cache. It’s cleaner, but requires a cache layer or library that supports it.

The write strategies

Write-through

Every write goes to the cache and the database at the same time, synchronously. Pros: the cache is always consistent with the database. Cons: slower writes, and data that is written but never read still fills the cache.

Write-back (write-behind)

Writes go to the cache first and are flushed to the database later, often in batches. Pros: very fast writes and fewer database operations. Cons: if the cache fails before flushing, data can be lost. Use it only where that risk is acceptable or mitigated.

Write-around

Writes go straight to the database, skipping the cache. The cache is populated only on reads. Pros: avoids filling the cache with data nobody reads. Cons: a read right after a write will miss the cache.

Comparison

StrategyRead latencyWrite latencyConsistencyRisk
Cache-asideFast after first readNormalCan be staleStale data
Read-throughFast after first readNormalCan be staleCache dependency
Write-throughFastSlowerStrongWasted cache space
Write-backFastFastestEventualData loss on failure
Write-aroundMiss after writesNormalGoodCold reads

Eviction policies

Caches are finite, so something has to go when they fill up:

  • LRU (least recently used): evicts what hasn’t been used for the longest time. A strong default.
  • LFU (least frequently used): evicts what is used least often. Good for stable popularity patterns.
  • FIFO: evicts the oldest entries. Simple, rarely optimal.
  • TTL-based expiry: entries expire after a set time, whatever their use.

Invalidation: the hard part

  • Set TTLs on everything, so stale data eventually disappears.
  • Delete or update keys on writes to the underlying data.
  • Version keys (for example, user:42:v7) so updates naturally create new entries.
  • Beware race conditions: a slow read can write stale data back into the cache just after an update clears it.

Protecting against cache stampedes

When a popular key expires, thousands of requests can miss at once and overwhelm the database. Defences include:

  • Request coalescing: only one request rebuilds the key; others wait.
  • Early refresh: refresh popular keys shortly before they expire.
  • Jittered TTLs: add randomness so many keys don’t expire at the same moment.

Where caches live

Browser caches, CDNs, reverse proxies, in-process memory, distributed caches such as Redis or Memcached, and database buffer caches all play a part. Many systems use several layers. When a distributed cache spans many nodes, consistent hashing is a common way to decide which node holds each key, so adding a node moves only a fraction of them.

Caching is also central to serving AI models efficiently: repeated prompts can be answered from cache rather than recomputed. AIEmulate’s guide to running AI models locally covers the hardware side.

Frequently asked questions

What is the most common caching strategy?

Cache-aside, because it’s simple and resilient: if the cache fails, the application can still read from the database. It’s also a safe default to propose in a system design interview.

How long should a cache TTL be?

As long as your data can safely be stale. Seconds for fast-changing data, hours or days for rarely changing data.

Should I cache everything?

No. Cache data that is read often, expensive to fetch and tolerant of brief staleness, such as the redirect lookups in a URL shortener.

Sources

  1. Redis documentation — caching patterns
  2. AWS — Caching best practices

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading