Caching Strategies Explained: Cache-Aside, Write-Through and Write-Back
Caching strategies and their trade-offs: cache-aside, read-through, write-through, write-back and write-around, plus eviction policies, TTLs and invalidation.
Rate limiting algorithms compared: fixed window, sliding log, sliding window counter, token bucket and leaky bucket, with code and distributed design tips.

Rate limiting controls how many requests a client can make in a period of time. It protects services from abuse and accidental overload, keeps usage fair between customers and controls costs. The algorithm you choose determines how smoothly limits are enforced and how much memory you need.
Count requests in fixed intervals, for example per minute. If the count exceeds the limit, reject until the next window.
Pros: very simple and memory-efficient. Cons: allows bursts at window boundaries. A client can send the full limit at the end of one minute and again at the start of the next, doubling the rate briefly.
Store a timestamp for every request. On each new request, drop timestamps older than the window and count what’s left.
Pros: precise, with no boundary bursts. Cons: memory grows with request volume, which is expensive at high traffic.
A compromise: combine the current and previous fixed windows, weighting the previous window by how much of it still overlaps the sliding window.
estimated = current_count + previous_count × (overlap fraction)
Pros: smooth limits with little memory. Cons: an approximation, though usually a good one.
A bucket holds up to capacity tokens and refills at a steady rate. Each request takes a token; if the bucket is empty, the request is rejected or delayed.
import time
class TokenBucket:
def __init__(self, capacity, refill_per_sec):
self.capacity = capacity
self.tokens = capacity
self.rate = refill_per_sec
self.updated = time.monotonic()
def allow(self):
now = time.monotonic()
self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)
self.updated = now
if self.tokens >= 1:
self.tokens -= 1
return True
return False
Pros: allows short bursts up to the bucket size while enforcing an average rate; memory-efficient. This is why it’s one of the most widely used algorithms. Cons: two parameters to tune.
Requests enter a queue that drains at a constant rate. When the queue is full, new requests are dropped.
Pros: produces a perfectly smooth outflow, useful for protecting fragile downstream systems. Cons: bursts are queued or dropped rather than served quickly, adding latency.
| Algorithm | Allows bursts | Memory | Accuracy | Typical use |
|---|---|---|---|---|
| Fixed window | Yes, at boundaries | Very low | Low | Simple quotas |
| Sliding log | No | High | Exact | Low-volume, strict limits |
| Sliding window counter | Limited | Low | Good | General API limits |
| Token bucket | Yes, controlled | Low | Good | APIs, gateways |
| Leaky bucket | No, smooths output | Low to medium | Good | Traffic shaping |
When a load balancer spreads requests across many servers, local counters aren’t enough. Common approaches:
AI APIs are a good real-world example: they often limit both requests and tokens per minute, which is one reason some teams choose to run models locally for heavy internal workloads.
Rate limiting is a favourite deep-dive topic in interviews; see our system design interview framework.
For most APIs, token bucket or sliding window counter. Choose leaky bucket when you need a perfectly smooth outflow.
Common keys are API key, user ID and IP address. Many systems combine several, with different limits for different endpoints.
Rate limiting rejects requests over a limit; throttling slows them down or queues them. The terms are often used interchangeably.
Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.
Caching strategies and their trade-offs: cache-aside, read-through, write-through, write-back and write-around, plus eviction policies, TTLs and invalidation.
Kafka, RabbitMQ and Amazon SQS compared: log vs queue models, ordering, replay, throughput, delivery guarantees and operations, and when to use each.
How idempotency keys make retries safe: why duplicate requests happen, how to store and check keys, handling concurrent requests, expiry and common mistakes.