Backend Patterns

Idempotency Keys: Designing Safe Retries

How idempotency keys make retries safe: why duplicate requests happen, how to store and check keys, handling concurrent requests, expiry and common mistakes.

A single glowing request passing through a gate while duplicates are blocked
Illustration: Backend Architect / AI-generated.

Key takeaways

  • Networks fail, clients retry and messages are redelivered, so duplicates are normal.
  • Clients send a unique key; the server stores the result and returns it for any repeat.
  • Handle concurrent duplicates with a lock or unique constraint, and expire keys after a sensible window.
On this page

A customer taps “Pay”. The request reaches your server and the payment succeeds, but the response is lost on a flaky mobile connection. The app retries. Without protection, you charge them twice. Idempotency keys prevent that.

What “idempotent” means

An operation is idempotent if doing it several times has the same effect as doing it once. HTTP GET, PUT and DELETE are designed to be idempotent; POST usually isn’t. Payments, orders and message sends need extra design to become safe to retry. Multi-service flows such as orders often use the saga pattern, where every step must be idempotent.

Why duplicates happen

  • Clients retry after timeouts, even though the first request succeeded.
  • Users double-click buttons.
  • Message queues deliver at least once, so consumers may see the same message twice; see Kafka vs RabbitMQ vs SQS.
  • Load balancers and proxies retry on errors.

How idempotency keys work

  1. The client generates a unique key (for example a UUID) for each logical operation and sends it in a header such as Idempotency-Key.
  2. The server checks whether it has seen the key before.
  3. New key: process the request, then store the key with the response.
  4. Seen key, finished: return the stored response without repeating the work.
  5. Seen key, still processing: return a conflict or wait, rather than processing again.
def create_payment(req, key):
    record = store.get(key)
    if record and record.done:
        return record.response
    if not store.claim(key):          # atomic insert; fails if another request holds the key
        return conflict("request in progress")
    response = charge_card(req)
    store.complete(key, response, ttl_hours=24)
    return response

Design details that matter

  • Atomic claim. Use a unique constraint or an atomic “set if not exists” so two concurrent duplicates can’t both proceed. The same trick underpins simple distributed locks.
  • Match the payload. If the same key arrives with a different request body, reject it: it’s a client bug.
  • Store the outcome, including errors that shouldn’t be retried.
  • Expire keys after a window long enough to cover realistic retries, commonly hours to a day.
  • Scope keys per client or account so they can’t collide across users.
  • Make side effects idempotent too. Downstream calls should carry the key or a derived ID.

Idempotent consumers

For queue consumers, record the IDs of processed messages, or design updates so repeating them is harmless, like setting a status rather than incrementing a counter.

Where it shows up in system design

Payments, order creation, notification sending and any API behind rate limits that clients retry with backoff.

Frequently asked questions

Who generates the idempotency key?

The client, once per logical operation, reusing it for every retry of that operation.

How long should keys be stored?

Long enough to cover retries and delayed redeliveries, typically 24 hours or so, depending on your system.

Is idempotency the same as exactly-once delivery?

No. Delivery may still happen more than once; idempotency makes the effect happen once.

Sources

  1. IETF — The Idempotency-Key HTTP Header Field (draft)
  2. MDN — Idempotent HTTP methods

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading