Series overview
Part 11 of 2839% complete
2026-06-20•22 min read

Saga: choreography and orchestration

Since Chapter 7, this series has carried a named, undefended gap: payment-service charges a customer on OrderPlaced without waiting to confirm inventory-service actually reserved stock. This chapter closes it, and shows both ways to do it.

1. Problem the Pattern Solves

Northwind finally hits the failure Chapter 7 predicted: during a flash sale, inventory-service finds a SKU is out of stock and publishes StockReservationFailed — but payment-service, listening independently to OrderPlaced on its own consumer group, has already charged the customer’s card by the time that failure event arrives. The order ends up FAILED in order-service’s eyes, while the customer has a real, captured charge on their statement for an order that will never ship. Support has to manually find and refund these — three of them last week alone.

The underlying problem is that “place an order” is not actually one operation — it’s a sequence of operations across three services (reserve stock, capture payment, confirm order) that must either all succeed or be undone in a coordinated way, but there is no distributed transaction spanning three independently-owned databases (Chapter 3 made sure of that, deliberately). Each service can only guarantee atomicity within its own boundary.

Forces in tension:

  • Consistency vs. availability. A saga achieves eventual consistency across services without locking them together in a distributed transaction — but “eventual” means there’s a real window where the system is in an intermediate, not-yet-reconciled state, exactly the window that caused Northwind’s refund problem.
  • Choreography’s simplicity vs. orchestration’s visibility. A choreographed saga (each service reacts to the previous step’s event and emits its own) has no single component to build or deploy — but the overall flow exists only implicitly, spread across services’ event handlers, hard to see as one thing. An orchestrated saga (a central coordinator issues commands and tracks state) makes the flow explicit and inspectable — at the cost of a new component that itself must be reliable and becomes a dependency in the flow.
  • Compensation complexity vs. correctness. Undoing a captured payment means issuing a real refund, not just “rolling back” — compensating actions are themselves business operations, with their own failure modes, not free actions the platform provides automatically.
  • Coupling. Choreography keeps services more loosely coupled to each other directly (each only cares about events, not who orchestrates), but couples all of them tightly to the shape and completeness of the event vocabulary — a poorly designed event schema makes for a fragile saga either way.

2. Core Idea

A saga is a sequence of local transactions across multiple services, coordinated so that either all steps complete successfully or the already-completed steps are undone by explicit compensating actions — since no cross-service ACID transaction is available (Chapter 3), the saga achieves the same business outcome atomicity would have provided, through coordination and compensation instead of locking.

Choreography: each service, on completing its local step, publishes an event; the next service(s) react to that event to perform their own step. No central coordinator — the saga’s logic is distributed across each participant’s event handlers.

payment-serviceinventory-serviceorder-servicepayment-serviceinventory-serviceorder-servicealt[payment succeeds][payment fails]alt[stock available][stock unavailable]create order (PENDING), publish OrderPlacedreserve stock(via Kafka) StockReservedcapture payment(via Kafka) PaymentCapturedmark PAID(via Kafka) PaymentFailedrelease stock (compensation)(via Kafka) PaymentFailedmark FAILED(via Kafka) StockReservationFailedmark FAILED
payment-serviceinventory-serviceorder-servicepayment-serviceinventory-serviceorder-servicealt[payment succeeds][payment fails]alt[stock available][stock unavailable]create order (PENDING), publish OrderPlacedreserve stock(via Kafka) StockReservedcapture payment(via Kafka) PaymentCapturedmark PAID(via Kafka) PaymentFailedrelease stock (compensation)(via Kafka) PaymentFailedmark FAILED(via Kafka) StockReservationFailedmark FAILED

Orchestration: a dedicated component (here, a saga orchestrator inside order-service) issues explicit commands to each participant in sequence and tracks the saga’s progress as its own piece of state, deciding what to do next — including which compensations to trigger — based on each participant’s response.

order-service (state)payment-serviceinventory-serviceOrderSagaOrchestratororder-service (state)payment-serviceinventory-serviceOrderSagaOrchestratoralt[payment fails][payment succeeds]create saga instance (state=STARTED)command: ReserveStockStockReservedupdate saga state=STOCK_RESERVEDcommand: CapturePaymentPaymentFailedcommand: ReleaseStock (compensation)StockReleasedsaga state=FAILEDPaymentCapturedsaga state=COMPLETED
order-service (state)payment-serviceinventory-serviceOrderSagaOrchestratororder-service (state)payment-serviceinventory-serviceOrderSagaOrchestratoralt[payment fails][payment succeeds]create saga instance (state=STARTED)command: ReserveStockStockReservedupdate saga state=STOCK_RESERVEDcommand: CapturePaymentPaymentFailedcommand: ReleaseStock (compensation)StockReleasedsaga state=FAILEDPaymentCapturedsaga state=COMPLETED

Commonly confused with:

  • A distributed transaction (2PC). A saga explicitly rejects holding locks across services for the duration of a multi-step operation — each local transaction commits independently and immediately; correctness comes from compensation after the fact, not from preventing other transactions from seeing intermediate state. This means a saga must tolerate other operations observing a not-yet-fully-committed-across-services state (Chapter 1’s AFTER_COMMIT discipline exists precisely because of this).
  • Workflow orchestration with BPMN (a later chapter). A BPMN engine typically models longer-running, often human-in-the-loop business processes with a graphical, business-analyst-editable definition. This chapter’s orchestrator is deliberately simpler — a plain Kotlin state machine — appropriate for a short-lived, fully-automated, developer-owned flow. The later chapter explains when the heavier BPMN tooling actually earns its place.
  • Retry with backoff (a later chapter). Retrying a single failed call is not a saga — a saga specifically addresses multi-step, cross-service operations where completed steps need explicit undoing, not just the failed step retried.

3. When to Use It

Strong indicators:

  • A business operation genuinely spans multiple services’ local transactions, and partial completion (some steps succeeded, later ones failed) is a real, financially or operationally meaningful problem — Northwind’s captured-payment-for-a-failed-order is exactly this.
  • No cross-service distributed transaction is available or desirable (true throughout this series since Chapter 3’s database-per-service migration).
  • Each step has, or can be given, a well-defined compensating action — if a step’s effect genuinely cannot be undone (an email already sent, in some interpretations), that step needs special handling (Section 7), not a saga pretending it’s reversible.

Concrete use cases:

  • E-commerce, as here: order placement spanning stock reservation and payment capture is the textbook case.
  • Travel booking: reserving a flight, a hotel, and a rental car across three (possibly external) systems, where any one failing must release the others — a saga is close to unavoidable here, since no single database spans all three providers.
  • Banking: a funds transfer between accounts held at different institutions (cross-bank, unlike the single-service transfer mentioned in Chapter 9) needs coordinated debit-and-credit steps with compensation if either leg fails.
  • Healthcare: scheduling a procedure that requires coordinated reservation of a room, a specialist’s calendar slot, and equipment — each owned by a different system, each needing release if any other step fails.

Prerequisites:

  • A complete, agreed inventory of every step’s compensating action, reviewed before implementation — an incomplete compensation strategy is worse than no saga, because it creates a false sense that failures are handled.
  • Idempotent commands and compensations (a saga step may be retried, and its compensation may itself need to be retried) — the idempotency discipline this series has required since Chapter 1 becomes non-negotiable here.
  • A decision, made deliberately per saga, between choreography and orchestration based on the criteria in Section 5 — not a default reached for out of habit.

4. When Not to Use It

  • The operation can be a single local transaction. If all the data a business operation touches lives in one service’s database, use a normal transaction (Chapter 1’s discipline) — a saga solves cross-service coordination, and applying it within one service’s boundary is needless complexity for a problem ACID already solves for free.
  • Strong consistency is a hard, non-negotiable requirement and the business can tolerate the operation being synchronous and slower. In rare cases, redesigning the boundary so the operation doesn’t cross services at all (or accepting a slower, more tightly-coupled synchronous chain with all its Chapter 2-documented costs) may genuinely be preferable to eventual consistency with compensation — this is a real trade-off to surface to the business, not assume away.
  • A step has no meaningful compensating action. If a step in the flow is truly irreversible (e.g., a physical action already taken in the world) and no reasonable compensation exists, a saga cannot fully solve the consistency problem — that step needs a different design (perhaps moved to be the last step in the sequence, after everything reversible has already succeeded).
  • Overengineering signal: building a full orchestrator with persistent saga state for a two-step flow where a simple choreographed pair of events, or even a synchronous call with a documented retry policy, would suffice. Match the mechanism’s weight to the flow’s actual length and failure-handling needs.

5. Implementation Example

Choreography, fixing the actual gap, by making payment-service wait for StockReserved rather than reacting to OrderPlaced directly — the minimal fix, using events this series already has in place since Chapters 7–8:

payment-service/src/main/kotlin/in/o612/eng/northwind/payment/internal/StockReservedListener.kt
package `in`.o612.eng.northwind.payment.internal
import `in`.o612.eng.northwind.inventory.api.StockReserved
import `in`.o612.eng.northwind.inventory.api.StockReservationFailed
import org.springframework.kafka.annotation.KafkaListener
import org.springframework.stereotype.Component
@Component
internal class StockReservedListener(private val paymentService: PaymentService) {
@KafkaListener(topics = ["inventory.events"], groupId = "payment-service")
fun onStockReserved(event: StockReserved) {
// Payment is now only ever attempted after stock is confirmed —
// the exact fix for Section 1's incident.
paymentService.captureForOrder(event.orderId)
}
@KafkaListener(topics = ["inventory.events"], groupId = "payment-service")
fun onStockReservationFailed(event: StockReservationFailed) {
// No compensation needed on payment's side — it never charged
// anything for this order, because it never received StockReserved.
}
}

The compensating action lives on inventory-service’s side, for the case where payment does fail after stock was reserved:

inventory-service/src/main/kotlin/in/o612/eng/northwind/inventory/internal/PaymentFailedListener.kt
package `in`.o612.eng.northwind.inventory.internal
import `in`.o612.eng.northwind.payment.api.PaymentFailed
import org.springframework.kafka.annotation.KafkaListener
import org.springframework.stereotype.Component
@Component
internal class PaymentFailedListener(private val inventoryService: InventoryService) {
@KafkaListener(topics = ["payment.events"], groupId = "inventory-service")
fun onPaymentFailed(event: PaymentFailed) {
// Compensating action: release the stock reserved for this order,
// since the saga cannot complete. Must be idempotent — safe to
// call twice if this event is redelivered.
inventoryService.releaseReservation(event.orderId)
}
}

This choreographed fix required no new component, no new deployable, and touches only the two services directly involved in the broken step — the strongest argument for choreography when the flow is short and the participants are already exchanging the right events.

Orchestration, for comparison — appropriate once the flow grows a third or fourth step and the implicit, spread-out logic of choreography becomes hard to see as one thing. This example adds a ScheduleDelivery step, making choreography’s “which service reacts to what” harder to follow at a glance, and introduces an explicit orchestrator instead:

order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/saga/OrderSagaOrchestrator.kt
package `in`.o612.eng.northwind.order.internal.saga
import org.springframework.stereotype.Component
import java.util.UUID
@Component
class OrderSagaOrchestrator(
private val sagaStateRepository: SagaStateRepository,
private val commands: SagaCommandGateway,
) {
fun start(orderId: UUID) {
sagaStateRepository.save(SagaState(orderId, SagaStep.STARTED))
commands.send(ReserveStockCommand(orderId))
}
fun onStockReserved(orderId: UUID) {
sagaStateRepository.updateStep(orderId, SagaStep.STOCK_RESERVED)
commands.send(CapturePaymentCommand(orderId))
}
fun onStockReservationFailed(orderId: UUID) {
sagaStateRepository.updateStep(orderId, SagaStep.FAILED)
// Nothing to compensate yet — this is the first step.
}
fun onPaymentCaptured(orderId: UUID) {
sagaStateRepository.updateStep(orderId, SagaStep.PAYMENT_CAPTURED)
commands.send(ScheduleDeliveryCommand(orderId))
}
fun onPaymentFailed(orderId: UUID) {
sagaStateRepository.updateStep(orderId, SagaStep.COMPENSATING)
commands.send(ReleaseStockCommand(orderId)) // explicit compensation, orchestrator-driven
}
fun onDeliveryScheduleFailed(orderId: UUID) {
sagaStateRepository.updateStep(orderId, SagaStep.COMPENSATING)
commands.send(RefundPaymentCommand(orderId)) // compensate step 2
commands.send(ReleaseStockCommand(orderId)) // compensate step 1
}
fun onSagaStepCompleted(orderId: UUID, step: SagaStep) {
if (step == SagaStep.DELIVERY_SCHEDULED) sagaStateRepository.updateStep(orderId, SagaStep.COMPLETED)
}
}
enum class SagaStep { STARTED, STOCK_RESERVED, PAYMENT_CAPTURED, DELIVERY_SCHEDULED, COMPLETED, COMPENSATING, FAILED }
data class SagaState(val orderId: UUID, val step: SagaStep)

Note what changed structurally, not just in line count: every transition and every compensation decision now lives in one file, readable top to bottom, with explicit persisted state (SagaState) recording exactly where each order’s saga is — directly answerable by a query, unlike choreography’s implicit state spread across which events have and haven’t fired yet.

order-service/src/test/kotlin/in/o612/eng/northwind/order/internal/saga/OrderSagaOrchestratorTest.kt
package `in`.o612.eng.northwind.order.internal.saga
import org.junit.jupiter.api.Test
import org.assertj.core.api.Assertions.assertThat
import java.util.UUID
class OrderSagaOrchestratorTest {
@Test
fun `payment failure triggers stock release, not delivery scheduling`() {
val orchestrator = OrderSagaOrchestrator(fakeStateRepo, spyCommandGateway)
val orderId = UUID.randomUUID()
orchestrator.start(orderId)
orchestrator.onStockReserved(orderId)
orchestrator.onPaymentFailed(orderId)
assertThat(spyCommandGateway.sentCommands).containsExactly(
ReserveStockCommand(orderId), CapturePaymentCommand(orderId), ReleaseStockCommand(orderId),
)
assertThat(fakeStateRepo.stateOf(orderId).step).isEqualTo(SagaStep.COMPENSATING)
}
}

6. Step-by-Step Flow

Tracing the choreographed fix (Section 5’s minimal version) for a payment failure after successful stock reservation:

  1. Client action. POST /orders, unchanged since Chapter 1.
  2. API request/event publication. order-service persists the order and publishes OrderPlaced, exactly as in Chapter 7.
  3. Service behavior. inventory-service reserves stock and publishes StockReserved. payment-service, now correctly waiting for this event rather than reacting to OrderPlaced directly, attempts to capture payment and it fails (card declined).
  4. Database interaction. payment-service records the failed attempt in its own schema; no charge is captured.
  5. Inter-service communication. payment-service publishes PaymentFailed. inventory-service, listening for it, releases the reservation — the compensating action.
  6. Error or failure handling. This is the error-handling path — the saga’s entire purpose. order-service, also listening for PaymentFailed, marks the order FAILED. No charge exists; no stock remains incorrectly reserved. Section 1’s incident cannot recur through this path.
  7. Observability signals. Track saga completion rate and compensation rate as explicit metrics — a rising compensation rate (more failed payments after successful reservations) is a meaningful business signal, not just a technical one, worth its own dashboard.
  8. Final response/outcome. The order ends in FAILED with no orphaned charge and no orphaned reservation — the correct, fully-compensated outcome that was previously only accidentally achieved, if at all.

7. Production Concerns

  • Timeouts and retries. Every saga step needs a timeout — if payment-service never responds (not even with a failure), the saga cannot progress or compensate correctly without one. Choreography needs each listener to have its own timeout/retry policy; orchestration can centralize this as a saga-level timeout that triggers compensation if a step doesn’t complete in time.
  • Idempotency, duplicate delivery. Every command and every compensation must be safe to execute more than once — releaseReservation for an already-released reservation should be a no-op, not an error, since Kafka’s at-least-once delivery (Chapter 7) guarantees redelivery will happen eventually.
  • Data consistency and transaction boundaries. The saga’s whole purpose is managing consistency without a cross-service transaction — but each individual step is still a proper local transaction within its own service, and the AFTER_COMMIT publish discipline from Chapters 1 and 7 remains essential at every step.
  • Compensation completeness. An incomplete compensation (e.g., releasing stock but forgetting to also refund a partially-processed payment in a longer chain) leaves the system in a worse, harder-to-detect state than no saga at all. Review every failure branch’s compensation chain explicitly, as the orchestrator example’s onDeliveryScheduleFailed does (compensating both prior steps, in reverse order).
  • API versioning. Saga command and event schemas need the same contract discipline as any other cross-service message (Chapters 2, 7, 8) — a saga adds more message types than a simple pub-sub flow, and each is a contract some other service depends on.
  • Authentication and service-to-service trust. An orchestrator issuing commands to multiple services needs credentials scoped appropriately for each — it’s effectively acting as a client to every participant service, and each participant should authorize commands as carefully as any other inbound request (Chapter 6’s concerns apply directly).
  • Logging, metrics, tracing, correlation IDs. A saga is the strongest case yet in this series for correlation-ID discipline — a multi-step, multi-service, asynchronous flow with compensation branches is genuinely difficult to debug without a trace that ties every step and every compensating action back to one originating order.
  • Kubernetes deployment. An orchestrator, if introduced, is a new deployable with its own scaling and availability needs — and, notably, it becomes a dependency every saga now runs through; plan its availability accordingly (multiple replicas, its own health checks) rather than treating it as an afterthought.
  • Testing strategy. Test every failure branch and its compensation explicitly (as Section 5’s orchestrator test does) — the success path is the easy 10% of a saga’s test surface; the failure and compensation paths are where correctness actually lives.
  • Migration strategy. Northwind fixed the two-step flow with choreography first (Section 5’s minimal fix), and only reaches for orchestration when a flow grows long or branchy enough that the implicit choreographed logic becomes hard to reason about — a decision revisited per flow, not decided once for the whole platform.

8. Common Mistakes

  1. Publishing a “success” event before the compensating actions for a possible later failure are even designed. Shipping the choreographed happy path without also implementing and testing every compensation branch — as Northwind’s original Chapter 7 flow did — leaves exactly the gap Section 1 describes. Fix: design and test compensations for every failure branch before considering a saga complete, not as a follow-up.
  2. Non-idempotent compensations. A releaseReservation that throws or double-releases stock when called twice (because the triggering event was redelivered) turns a safety mechanism into a new bug source. Fix: every compensation must be idempotent, verified by an explicit test that calls it twice.
  3. Choosing orchestration by default “because it’s more visible,” even for a two-step flow. Building a full orchestrator with persisted saga state for Northwind’s original two-step reserve-then-charge flow is more machinery than the flow’s complexity warrants. Fix: start with choreography for short, simple flows; introduce orchestration when the flow’s length or branching genuinely outgrows choreography’s implicit-logic readability, as Section 5 demonstrates with the three-step delivery-scheduling example.
  4. Treating a saga as if it provides isolation. Assuming no other process can observe an order in its “stock reserved, payment not yet captured” intermediate state is false — a saga explicitly does not provide the isolation a distributed transaction would have. Fix: design any code that reads order state (dashboards, customer-facing status pages) to handle and correctly represent intermediate saga states, not just terminal ones.
  5. No timeout on a saga step. A step that can hang forever (a downstream payment processor that never responds) leaves the saga stuck indefinitely, with stock reserved and no resolution. Fix: every step needs an explicit timeout that triggers a defined outcome — proceed, retry, or begin compensation — never silence.
  6. Letting the orchestrator become a second source of business logic that drifts from what each service actually enforces. If the orchestrator’s understanding of “when can an order be cancelled” diverges from order-service’s own domain rules (Chapter 9’s event-sourced validation, for instance), the two can produce contradictory outcomes. Fix: the orchestrator coordinates sequencing; each service remains the sole authority over its own business rules — the orchestrator should never duplicate or second-guess them.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Multi-step operation spans services, partial completion is a real (financial/operational) problemYesSaga provides coordinated, compensatable eventual consistency where no distributed transaction is available—
Flow is short (2-3 steps), participants already exchange the right eventsChoreographyNo new component; simplest fix for the coupling already in place—
Flow is long, branchy, or hard to reason about as scattered event handlersOrchestrationCentralizes and makes explicit the sequencing and compensation logic—
Operation fits within a single service’s transaction boundaryNoA saga solves a cross-service problem you don’t haveLocal ACID transaction
A step is genuinely irreversible with no workable compensationReconsider the flow’s ordering, or accept manual intervention for that stepA saga can’t fully automate around a truly non-compensatable actionReorder so irreversible steps happen last, after all reversible steps succeed

10. Hands-On Exercise

Extend it: add a ScheduleDelivery step to the choreographed version from Section 5 (not just the orchestrated example), and observe firsthand how much harder it becomes to answer “what exactly happens if delivery scheduling fails after payment succeeded” by reading scattered event listeners, compared to the orchestrator’s explicit onDeliveryScheduleFailed method. Use that experience to justify, concretely, when you’d switch.

Simulate a failure: in the orchestrated version, make ScheduleDelivery fail after CapturePayment has already succeeded. Confirm both compensations fire (refund the payment, release the stock) in the correct order, and that the saga’s persisted state correctly reflects COMPENSATING then FAILED.

Decision question, with justification required: Northwind’s saga now needs to also notify the customer by email at each terminal state (COMPLETED or FAILED). Should the notification step be part of the saga itself (with its own compensation if it fails — arguably impossible, since an email can’t be unsent) or handled entirely outside the saga, as a simple, best-effort reaction to the saga’s terminal event? Justify using Section 4’s “no meaningful compensation” criterion.

11. Key Takeaways

  • A saga coordinates a multi-step, cross-service business operation into eventual consistency through local transactions plus explicit compensating actions — it does not, and cannot, provide the isolation or atomicity a distributed transaction would.
  • Choreography (event reactions) fits short, simple flows with minimal new infrastructure; orchestration (an explicit coordinator with persisted state) fits longer or branchier flows where implicit, scattered logic becomes hard to reason about — choose per flow, not once for the whole platform.
  • Every step needs a genuine, idempotent compensating action designed and tested before the saga is considered complete — the happy path is the easy part.
  • A saga explicitly does not provide isolation — other parts of the system can and will observe intermediate states, and must be designed to handle that correctly.
  • Timeouts on every step are mandatory — a step with no timeout can leave a saga stuck indefinitely in a partially-completed state.
  • An orchestrator coordinates sequencing; it must never duplicate or override a participant service’s own business rules, or the two can produce contradictory outcomes.
  • This pattern exists specifically because Chapter 3 ruled out cross-service distributed transactions — it’s the direct, necessary consequence of database-per-service, not an optional add-on.
Spring BootKotlinMicroservicesKafka

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind