Distributed Systems

The Saga Pattern for Distributed Transactions

The saga pattern explained: local transactions and compensating actions, choreography vs orchestration, the outbox pattern and isolation pitfalls.

A long winding line of black dominoes, the first few toppled and the rest still standing
Illustration: Backend Architect / AI-generated.

Key takeaways

  • A saga replaces one distributed transaction with a sequence of local transactions, each with a compensating action.
  • Choreography links steps through events; orchestration uses a central coordinator that is easier to follow and test.
  • Sagas aren’t isolated: design for intermediate states, make every step idempotent and publish events reliably with an outbox.
On this page

The saga pattern manages a business transaction that spans several services, each with its own database, without a distributed lock or two-phase commit. The process is split into a sequence of local transactions. Each step commits in its own service and triggers the next. If a step fails, the saga runs compensating actions that undo the earlier steps in business terms, such as refunding a payment or releasing reserved stock. The result is eventual consistency instead of one atomic commit.

Why not a normal transaction?

Inside one database, an ACID transaction makes a multi-step change atomic; our guide to transaction isolation levels explains what it guarantees. Across services, a single transaction would need two-phase commit (2PC), which holds locks while waiting on every participant, blocks if the coordinator fails and is often unsupported by message brokers and modern databases. Sagas, first described by Garcia-Molina and Salem in 1987 for long-lived database transactions, avoid holding locks across the whole process.

An example: placing an order

  1. Order service: create the order as PENDING.
  2. Payment service: charge the card.
  3. Inventory service: reserve the items.
  4. Shipping service: schedule delivery.
  5. Order service: mark the order CONFIRMED.

If inventory can’t reserve the items at step 3, the saga compensates: refund the payment (undoing step 2) and mark the order CANCELLED (undoing step 1).

StepActionCompensation
1Create order (pending)Cancel order
2Charge paymentRefund payment
3Reserve inventoryRelease reservation
4Schedule shippingCancel shipment

Compensations are business operations, not database rollbacks: a refund is a new transaction that leaves an audit trail, which is exactly what finance teams want.

Choreography vs orchestration

Choreography: each service listens for events and reacts. The order service emits OrderCreated; payment hears it, charges the card and emits PaymentCompleted; inventory hears that, and so on.

  • Loose coupling and no central component.
  • Hard to see the whole flow; logic is scattered, and cyclic dependencies creep in as the saga grows.

Orchestration: a saga orchestrator tells each service what to do and tracks the state, stepping forward or compensating backward.

  • The flow is explicit in one place, easier to test, monitor and change.
  • The orchestrator is another component to build and run, and it must persist its state.
ChoreographyOrchestration
ControlDistributed, via eventsCentral coordinator
VisibilityLowHigh
CouplingServices know each other’s eventsServices know the orchestrator’s commands
Best forShort sagas, 2–4 stepsLonger or frequently changing flows

Workflow engines such as Temporal, AWS Step Functions and Camunda provide durable orchestration so you don’t have to build state machines by hand.

Reliable messaging: the outbox pattern

A step usually has to update its database and publish an event. Doing those separately risks one succeeding without the other. The transactional outbox pattern writes the event to an outbox table in the same local transaction as the business change; a relay then publishes outbox rows to the broker. Our comparison of Kafka, RabbitMQ and SQS covers the broker side. Because relays and brokers deliver at least once, every handler must be idempotent; see idempotency keys.

The isolation problem

Sagas give up isolation: other requests can see intermediate states, such as an order that is charged but not yet confirmed. Common countermeasures:

  • Semantic locks: mark records as pending so other operations treat them carefully.
  • Commutative updates that give the same result in any order.
  • Reread before acting: verify data hasn’t changed before a critical step.
  • Order steps wisely: put steps that can’t be compensated, such as sending goods, after the point of no return, and make them retriable until they succeed.

Designing good sagas

  • Model states explicitly: pending, confirmed, cancelled, compensating.
  • Make compensations idempotent and retriable; they can fail too.
  • Set timeouts for steps that never answer, then compensate or escalate.
  • Log a correlation ID across all steps for tracing.
  • Keep sagas short. A long chain of services is a hint that boundaries may be wrong.

Testing and observing sagas

Test the unhappy paths deliberately: make each step fail in turn and check that compensations leave every service consistent. In production, track how many sagas are in progress, how long each step takes and how often compensations run. A sudden rise in compensations is often the first sign that a dependency is failing.

When to use sagas

Use them when a business process genuinely spans services that own their own data. If every step touches the same database, a local transaction is simpler and safer. Sagas pair naturally with event sourcing and CQRS, and the need for them is one of the real costs weighed in our microservices vs monolith framework.

Frequently asked questions

Is a saga the same as two-phase commit?

No. Two-phase commit makes all participants commit atomically while holding locks. A saga commits each step immediately and uses compensations to undo them if something fails later.

What happens if a compensation fails?

Retry it, since compensations should be idempotent. If it keeps failing, alert a human and park the saga in a visible failed state; never drop it silently.

Choreography or orchestration?

Choreography suits short, stable flows. Orchestration is easier to understand and change once a saga has more than a few steps.

Sources

  1. Garcia-Molina and Salem — Sagas (1987)
  2. Chris Richardson — Pattern: Saga
  3. Chris Richardson — Pattern: Transactional outbox

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading

Distributed Systems

How CDNs Work

How CDNs work: edge locations, request routing, cache keys, Cache-Control headers, invalidation, origin shielding, security features and edge compute.

4 min read