Distributed Systems

Distributed Locks and Leases

How distributed locks and leases work, why process pauses break naive locks, fencing tokens, Redis, ZooKeeper, etcd and database locks, and alternatives.

A row of identical steel padlocks on a metal rail, one of them open, in dramatic side light
Illustration: Backend Architect / AI-generated.

Key takeaways

  • A distributed lock should be a lease with an expiry, because the holder can crash or stall.
  • Leases alone are unsafe for correctness: a paused process can wake up after its lease expired; fencing tokens fix this.
  • Use a consensus-backed service or your database for correctness-critical locks, and avoid locks entirely where idempotency or partitioning will do.
On this page

A distributed lock lets one process at a time, across many machines, do something that must not happen concurrently: run a scheduled job, process a file, or update a shared resource. Because the holder can crash, locks are granted as leases that expire. The hard part is that a process can pause (for garbage collection, a slow disk or a network stall), outlive its lease and carry on as if it still held it. For correctness, pair leases with fencing tokens, and prefer a consensus-backed store or your database over ad-hoc schemes.

First, ask what the lock is for

Kleppmann’s widely cited analysis splits locks into two kinds:

  • Efficiency locks stop duplicate work, such as two servers sending the same report. If the lock occasionally fails, you waste some effort or send a duplicate email. A simple lock is fine.
  • Correctness locks protect data. If two holders act at once, data is corrupted or money moves twice. These need much stronger guarantees.

Be honest about which you have, because it decides the design.

Leases: locks that expire

A lease is a lock with a time limit. If the holder crashes, the lease runs out and someone else can take over, so the system never deadlocks on a dead process. Long-running holders renew the lease periodically.

The catch: the holder can’t be sure its lease is still valid. A 30-second garbage-collection pause, a VM migration or a slow network can mean it resumes after expiry, while another process already holds the lock. Checking the clock just before writing doesn’t help, because the pause can happen right after the check.

Fencing tokens

The robust fix is a fencing token: every time the lock is granted, the lock service hands out a number that only ever increases. The holder sends the token with every write, and the storage system rejects any write carrying a token lower than one it has already seen.

StepClient AClient BStorage
1Gets lock, token 33
2Pauses (GC)
3Lease expiresGets lock, token 34
4Writes with token 34Accepts; highest seen is 34
5Wakes, writes with token 33Rejects: 33 is lower than 34

Google’s Chubby lock service described a similar idea, called sequencers. Fencing requires the protected resource to check tokens, which is easy for a database row with a version column and harder for third-party APIs.

Implementation options

Your database

If the protected data lives in one relational database, use it: row locks with SELECT … FOR UPDATE, advisory locks such as PostgreSQL’s pg_try_advisory_lock, or a lock table with a unique constraint and an expiry column. Locks and data then share one source of truth. Our guide to transaction isolation levels covers row locking and optimistic concurrency, which often removes the need for a separate lock.

Redis

For efficiency locks, a single Redis instance works well. Set a key only if it doesn’t exist, with an expiry and a unique random value:

SET lock:nightly-report token-8f3a NX PX 30000

Release it only if the value still matches, using a small script, so you never delete a lock someone else now holds. Redis also documents Redlock, an algorithm across several independent Redis nodes; its safety depends on timing assumptions that Kleppmann and others have challenged, so don’t rely on it where correctness is at stake without fencing.

ZooKeeper and etcd

These are consensus-based coordination services designed for exactly this problem. Locks are tied to a client session or lease: ZooKeeper uses ephemeral sequential nodes that vanish if the client’s session dies, and etcd attaches keys to leases. Both expose ever-increasing revision or version numbers that can serve as fencing tokens, and both stay consistent through node failures, at the cost of running a quorum cluster.

Comparing the options

OptionSafetyOperational costGood for
Database locksStrong, within one databaseLow (you already run it)Protecting that database’s data
Single Redis instanceWeak: fails with the instanceLowEfficiency locks, deduplicating jobs
RedlockDebated; timing-dependentMediumEfficiency locks across nodes
ZooKeeper / etcdStrong (consensus)Medium to highLeader election, correctness locks

Often you don’t need a lock

  • Make the operation idempotent so doing it twice is harmless; see idempotency keys.
  • Use compare-and-set or version checks at the data store instead of a separate lock.
  • Give each piece of work a single owner by partitioning: route all updates for a key to one worker or shard, as in database sharding.
  • Use a queue with one consumer per partition, which serialises work without locks.

Practical rules

  • Always set an expiry; never hold a lock without one.
  • Choose lease times that comfortably exceed normal work time, and renew for long jobs.
  • Release only your own lock, checked by a unique value.
  • Use fencing tokens wherever a stale holder could corrupt data.
  • Expect lock-service unavailability, and decide whether to fail closed or open. Network partitions force exactly the trade-off described by the CAP theorem and PACELC.

Frequently asked questions

Is a Redis lock safe?

For efficiency, usually yes. For correctness, a single instance can lose the lock on failover, and pauses can outlive leases, so add fencing tokens or use a consensus-based system.

What is the difference between a lock and a lease?

A lease is a lock with an expiry. In distributed systems, locks should almost always be leases, so a crashed holder can’t block everyone for ever.

How long should a lease be?

Long enough to cover normal processing with a margin, and short enough that recovery after a crash is acceptable. Renew it for longer jobs rather than setting a huge timeout.

Sources

  1. Martin Kleppmann — How to do distributed locking (2016)
  2. Redis — Distributed locks with Redis
  3. Burrows — The Chubby lock service for loosely-coupled distributed systems (2006)

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading

Distributed Systems

How CDNs Work

How CDNs work: edge locations, request routing, cache keys, Cache-Control headers, invalidation, origin shielding, security features and edge compute.

4 min read