Backend Patterns

Load Balancing Algorithms Explained

Load balancing algorithms compared: round robin, weighted, least connections, least response time, IP hash and consistent hashing, L4 vs L7 and health checks.

Traffic streams being evenly distributed across several servers
Illustration: Backend Architect / AI-generated.

Key takeaways

  • Round robin is simple and works well when servers and requests are similar.
  • Least connections and least response time adapt to uneven request costs.
  • Hash-based algorithms give session affinity; health checks remove failed servers.
On this page

A load balancer spreads incoming requests across multiple servers so no single machine is overwhelmed, and so the service keeps running when one fails. The algorithm it uses to pick a server shapes latency, fairness and resilience.

The main algorithms

AlgorithmHow it picks a serverBest for
Round robinEach server in turnSimilar servers, similar requests
Weighted round robinIn turn, weighted by capacityMixed server sizes
Least connectionsFewest active connectionsLong or uneven requests
Least response timeFastest recent responses and fewest connectionsLatency-sensitive services
Random (power of two choices)Pick two at random, choose the less loadedLarge fleets, simple and effective
IP or header hashHash of client IP or a headerSession affinity
Consistent hashingHash ring of serversCaches and stateful services

Round robin and weighted round robin

The simplest approach, and often good enough. Weighting lets bigger servers take more traffic. Its weakness: it ignores how busy each server actually is.

Least connections and least response time

These adapt to reality: if some requests are slow (file uploads, reports), servers handling them receive fewer new requests. They need the balancer to track state per server.

Power of two choices

Picking two servers at random and sending the request to the less loaded one performs remarkably close to “always pick the best”, with much less coordination. It’s popular in large distributed systems.

Hashing and session affinity

Hash-based algorithms send the same client to the same server, useful when servers hold session state or caches. Consistent hashing keeps most assignments stable when servers are added or removed. Better still, keep servers stateless and store sessions in a shared cache; see caching strategies.

Layer 4 vs layer 7

  • Layer 4 (transport): routes by IP and port; very fast, but can’t see HTTP details.
  • Layer 7 (application): routes by URL, headers or cookies; enables path-based routing, TLS termination and smarter decisions.

Health checks and failover

Balancers regularly probe servers and stop sending traffic to unhealthy ones. Combine health checks with timeouts, retries (made safe by idempotency keys) and connection draining during deployments.

In system design

Place load balancers in front of every stateless tier, and mention algorithm choice when request costs vary. For global systems, add DNS or anycast-based balancing across regions. CDNs use the same techniques to send users to a nearby edge; see how CDNs work.

Frequently asked questions

Which load balancing algorithm is best?

There’s no single best. Round robin is fine for uniform workloads; least connections or response time suit uneven ones.

What are sticky sessions?

Sending a client to the same server each time, usually via a cookie or hash. Useful for stateful apps, but it can unbalance load.

Is a load balancer a single point of failure?

It can be. Production setups run balancers in redundant pairs or use managed, distributed balancers.

Sources

  1. NGINX — HTTP load balancing
  2. HAProxy documentation

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading