Event sourcing
Every service so far in this series — order-service included — stores its current state as mutable rows (Chapter 1’s order_schema, unchanged in shape since). This chapter asks whether order-service should keep doing that, and answers: for one specific reason, yes to the log, no to the mutable row, and explains exactly why that reason doesn’t generalize to inventory-service or payment-service.
1. Problem the Pattern Solves
A customer disputes a charge, claiming their order was never actually cancelled despite Northwind’s records showing status = CANCELLED. Support pulls up the order row: it shows the current status and a single updatedAt timestamp. Nothing in the database says who cancelled it, when each prior status change happened, or what the order looked like at the moment the customer says they confirmed it. The UPDATE orders SET status = ? WHERE id = ? that Chapter 1’s OrderStatusUpdater runs on every transition (PENDING → AWAITING_PAYMENT → PAID/FAILED/CANCELLED) destroys the previous value every time — the row remembers only where the order is, never how it got there.
Separately, Northwind’s finance team wants a report answering “what was our total reserved-but-unpaid order value at the close of business on any given day last quarter” — a question the current-state table cannot answer at all, because it was never designed to represent history, only the present.
Forces in tension:
- Auditability vs. storage and query simplicity. A mutable current-state row is compact and trivially queryable (“give me order X’s status”) but destroys history by construction. An append-only event log preserves complete history but requires replaying events to answer even the simplest “what is order X’s status right now” question, unless you also maintain a projection.
- Debuggability vs. complexity. Being able to replay exactly the sequence of events that led to a bug is powerful for diagnosis — but the team now has two things to keep consistent (the event log and any derived read model), where before there was only one table.
- Flexibility of new views vs. up-front modeling cost. New questions about history (“what did this look like at time T”) become answerable later, from the same log, without a schema migration — but only if the events captured enough detail up front, which requires more careful domain modeling than “just add a column” does.
- Team familiarity vs. genuine benefit. Event sourcing is a more specialized, less commonly practiced pattern than CRUD persistence — the learning curve and the risk of getting the event schema wrong are real costs that must be weighed against a genuine auditability or temporal-query requirement, not general appeal.
2. Core Idea
Event sourcing means storing a service’s state as an ordered, append-only sequence of domain events (OrderCreated, OrderStockReserved, OrderPaymentCaptured, OrderCancelled) rather than as a mutable “current state” row. The current state of any entity is derived by replaying its events from the beginning (or from the last snapshot — see Section 5), never stored as the sole source of truth.
Intent: make the history of how an entity reached its current state a first-class, permanently available fact, not a side effect that gets destroyed by the next UPDATE.
Participants:
- Event store — the append-only log itself; here, a Postgres table (
order_events), since Northwind’s volume doesn’t yet justify a specialized event-store database, and Postgres already has the operational maturity the team trusts (Chapter 3). - Aggregate — the entity whose full history is sourced (
Order), rebuilt by folding its events in sequence. - Projection — a derived, queryable view built from the event log (the familiar
order_summarycurrent-state table), kept for read performance so that “what’s order X’s status” doesn’t require replaying its whole history on every query.
Commonly confused with:
- Event-driven architecture (Chapter 7). Event-driven architecture is about how services communicate with each other; event sourcing is about how a service persists its own internal state.
order-servicepublishingOrderPlacedto Kafka for other services (Chapter 7) is unrelated to whetherorder-serviceitself stores its own history as an event log — this chapter adds the latter without changing the former. - CQRS (next chapter). Event sourcing and CQRS are frequently paired but are independent decisions: you can event-source a service’s writes while still reading from the same normalized-ish tables (less common), or use CQRS’s read/write separation over an ordinary mutable-row write model (very common, and often sufficient on its own). This chapter’s
order_summaryprojection is, in fact, the seed of the next chapter’s full CQRS treatment. - An audit log bolted onto a CRUD table. A separate
order_audit_logtable populated by a trigger or an@PostUpdatehook alongside the “real” mutable table is a legitimate, much simpler alternative for auditability alone (see Section 4) — it is not event sourcing, because the audit log isn’t the source of truth the current state is derived from; the mutable table still is.
3. When to Use It
Strong indicators:
- A genuine, named business or compliance requirement to know how an entity reached its current state, not just what the current state is — Northwind’s dispute-resolution and “state at a point in time” reporting needs are exactly this.
- The domain has meaningful, named business events (not just column updates) — “stock reserved,” “payment captured,” “order cancelled” are events a domain expert would recognize and name; a generic
updated_attimestamp on a CRUD row is not. - The entity’s history itself has value beyond debugging — e.g., replaying events to answer new business questions later without a schema migration.
Concrete use cases:
- Order management, as here: dispute resolution, “what did this order look like when the customer confirmed it,” and point-in-time financial reporting all require history a mutable row destroys.
- Banking and payments ledgers: a transaction history is, definitionally, an append-only sequence of events — double-entry bookkeeping is essentially event sourcing that predates the software pattern by centuries.
- Healthcare records: a patient’s clinical history must be auditable and never silently overwritten — regulatory requirements often mandate exactly the “why did this change, who changed it, when” guarantees event sourcing provides natively.
- Insurance claims processing: a claim’s full lifecycle (filed, reviewed, disputed, adjusted, settled) is inherently a sequence of meaningful state transitions that adjusters and auditors need to review in full, not just the final outcome.
Prerequisites:
- Careful up-front domain modeling of what events actually mean and what fields they carry — an event schema is even more consequential to get right than a database schema, since past events are typically never rewritten, only ever appended to or superseded by new event types.
- A plan for rebuilding projections when a new read need arises, and an upcasting/versioning strategy for evolving event schemas over time as the domain model inevitably changes (Section 7).
- Acceptance that most queries will go through a projection, not the raw event log — treating the log itself as a queryable table for arbitrary questions defeats its purpose and performs poorly.
4. When Not to Use It
- No genuine history requirement.
inventory-service’s stock count doesn’t need its full mutation history for any named business reason Northwind has today — a plain mutable row (Chapter 1’sStockRepository, unchanged) remains the right choice there. Adopting event sourcing platform-wide “for consistency” withorder-servicewould add real complexity to two services with no corresponding benefit. - A simple audit trail is enough. If the only requirement is “know who changed what and when,” a much simpler audit-log table or Hibernate Envers-style history table, alongside an ordinary mutable current-state table, satisfies it without the replay/projection machinery this pattern requires.
- Team unfamiliarity with the pattern under time pressure. Event sourcing’s learning curve (aggregate reconstruction, snapshotting, projection rebuilding, schema evolution) is real; adopting it for the first time under a tight deadline, on a service where a mutable row would have sufficed, multiplies delivery risk for a benefit the team hasn’t yet proven it needs.
- Overengineering signal: applying event sourcing to every entity in a domain because one entity (here,
Order) genuinely benefited from it.inventory-service’s stock level andpayment-service’s charge record don’t shareOrder’s auditability requirement — evaluate each aggregate independently, exactly as Chapter 3 evaluated each service’s database needs independently.
5. Implementation Example
Only order-service adopts this pattern in Northwind’s system — inventory-service and payment-service keep their Chapter 1 mutable-row persistence entirely unchanged. This is the point: event sourcing is an internal implementation choice of one service, invisible to its callers, exactly like database-per-service (Chapter 3) was.
The event log, an append-only table — never updated, only inserted into:
package `in`.o612.eng.northwind.order.internal.eventsourcing
import java.math.BigDecimalimport java.time.Instantimport java.util.UUID
sealed interface OrderEvent { val orderId: UUID val occurredAt: Instant
data class OrderCreated( override val orderId: UUID, override val occurredAt: Instant, val customerId: UUID, val items: List<LineItem>, ) : OrderEvent
data class StockReservationConfirmed( override val orderId: UUID, override val occurredAt: Instant, ) : OrderEvent
data class StockReservationFailed( override val orderId: UUID, override val occurredAt: Instant, val unavailableSkus: List<String>, ) : OrderEvent
data class PaymentCaptured( override val orderId: UUID, override val occurredAt: Instant, val amount: BigDecimal, ) : OrderEvent
data class OrderCancelled( override val orderId: UUID, override val occurredAt: Instant, val reason: String, val cancelledBy: String, ) : OrderEvent}
data class LineItem(val sku: String, val quantity: Int, val unitPrice: BigDecimal)package `in`.o612.eng.northwind.order.internal.eventsourcing
import org.springframework.jdbc.core.JdbcTemplateimport org.springframework.stereotype.Repositoryimport java.util.UUID
@Repositoryinternal class OrderEventStore(private val jdbc: JdbcTemplate, private val serializer: OrderEventSerializer) {
/** Insert-only — this table never receives an UPDATE or DELETE in normal * operation. `sequence_number` enforces per-aggregate ordering and * doubles as an optimistic-concurrency check on append. */ fun append(orderId: UUID, expectedSequence: Long, event: OrderEvent) { val rows = jdbc.update( """INSERT INTO order_events (order_id, sequence_number, event_type, payload, occurred_at) SELECT ?, ?, ?, ?::jsonb, ? WHERE (SELECT COUNT(*) FROM order_events WHERE order_id = ?) = ?""", orderId, expectedSequence + 1, event::class.simpleName, serializer.toJson(event), event.occurredAt, orderId, expectedSequence, ) if (rows == 0) throw ConcurrentModificationException("Order $orderId was modified concurrently") }
fun loadEvents(orderId: UUID): List<OrderEvent> = jdbc.query("SELECT payload, event_type FROM order_events WHERE order_id = ? ORDER BY sequence_number", orderId) { rs, _ -> serializer.fromJson(rs.getString("payload"), rs.getString("event_type")) }}CREATE TABLE order_events ( order_id UUID NOT NULL, sequence_number BIGINT NOT NULL, event_type TEXT NOT NULL, payload JSONB NOT NULL, occurred_at TIMESTAMPTZ NOT NULL, PRIMARY KEY (order_id, sequence_number));The WHERE (SELECT COUNT(*) ... ) = ? guard is a simple optimistic-concurrency check: appending event N+1 only succeeds if exactly N events already exist for that order, preventing two concurrent commands from silently producing conflicting histories.
Rebuilding an Order aggregate by folding its events — this replaces Chapter 1’s OrderRepository.findById for any code path that needs the current business state to make a decision, such as validating a cancellation is allowed:
package `in`.o612.eng.northwind.order.internal.eventsourcing
data class Order( val id: java.util.UUID, val status: OrderStatus, val items: List<LineItem>, val sequenceNumber: Long,) { companion object { fun rebuild(events: List<OrderEvent>): Order = events.fold(EMPTY) { order, event -> order.apply(event) } .copy(sequenceNumber = events.size.toLong())
private val EMPTY = Order(java.util.UUID(0, 0), OrderStatus.NONE, emptyList(), 0) }
private fun apply(event: OrderEvent): Order = when (event) { is OrderEvent.OrderCreated -> copy(id = event.orderId, status = OrderStatus.PENDING, items = event.items) is OrderEvent.StockReservationConfirmed -> copy(status = OrderStatus.AWAITING_PAYMENT) is OrderEvent.StockReservationFailed -> copy(status = OrderStatus.FAILED) is OrderEvent.PaymentCaptured -> copy(status = OrderStatus.PAID) is OrderEvent.OrderCancelled -> copy(status = OrderStatus.CANCELLED) }}
enum class OrderStatus { NONE, PENDING, AWAITING_PAYMENT, PAID, FAILED, CANCELLED }package `in`.o612.eng.northwind.order.internal.eventsourcing
import java.time.Instantimport java.util.UUID
class CancelOrderHandler(private val eventStore: OrderEventStore) {
fun cancel(orderId: UUID, reason: String, cancelledBy: String) { val order = Order.rebuild(eventStore.loadEvents(orderId)) require(order.status in setOf(OrderStatus.PENDING, OrderStatus.AWAITING_PAYMENT)) { "Cannot cancel an order in status ${order.status}" } eventStore.append(orderId, order.sequenceNumber, OrderEvent.OrderCancelled(orderId, Instant.now(), reason, cancelledBy)) }}Notice the business rule — an order can only be cancelled while PENDING or AWAITING_PAYMENT — is enforced against the replayed state, derived fresh from the full event history every time, never from a value that could have been silently overwritten.
The projection, kept for fast reads — this is what GET /orders/{id} actually queries, never the raw event log directly:
package `in`.o612.eng.northwind.order.internal.eventsourcing
import org.springframework.stereotype.Component
/** Rebuilt from the event log after every append; kept in an ordinary * `order_summary` table with the same shape Chapters 1-8 already used, so * every existing REST/gRPC/event contract is unaffected by this migration. */@Componentclass OrderSummaryProjector(private val summaryRepository: OrderSummaryRepository) {
fun project(orderId: UUID, events: List<OrderEvent>) { val order = Order.rebuild(events) summaryRepository.upsert(OrderSummaryRow(order.id, order.status.name, order.items.sumOf { it.unitPrice.multiply(it.quantity.toBigDecimal()) })) }}6. Step-by-Step Flow
- Client action.
POST /orders/{id}/cancel— a new endpoint this pattern makes possible to implement correctly, since the cancellation rule needs to know the order’s actual history-derived state, not just a possibly-stale current-status column. - API request.
order-service’s handler loads and replays the order’s full event history. - Service behavior. The business rule (only
PENDING/AWAITING_PAYMENTorders can be cancelled) is evaluated against the freshly rebuilt state. - Database interaction. A new
OrderCancelledevent is appended (never an update) with an optimistic-concurrency check against the expected sequence number. - Inter-service communication. After the append succeeds,
order-servicestill publishesOrderCancelledto Kafka exactly as Chapter 7 established — event sourcing (internal persistence) and event-driven architecture (inter-service communication) coexist and reinforce each other here, the domain event doing double duty as both the store’s record and the integration event, a natural and common pairing. - Error or failure handling. A concurrent modification (someone else appended an event for the same order between load and append) throws and the client retries — a real, structural guarantee against lost updates that Chapter 1’s plain
UPDATE ... WHERE id = ?never provided. - Observability signals. Track event-append rate and projection-rebuild lag per aggregate type — if the projection updater falls behind the event log, reads become stale in a way that’s invisible unless specifically monitored.
- Final response. The client gets
200 OK; the order’s full history — including the now-supersededAWAITING_PAYMENTstate — remains permanently queryable for the dispute-resolution and reporting needs from Section 1.
7. Production Concerns
- Timeouts, retries, idempotency. Appending an event must be idempotent under retry — use a client-supplied idempotency key or the optimistic-concurrency check itself (a retried append with a stale expected-sequence number simply fails safely, as shown above) rather than relying on the caller to avoid double-submission.
- Data consistency and transaction boundaries. The event append and the projection update should ideally happen in the same transaction (as in Section 5’s example, both against the same Postgres instance) — if they can’t be (e.g., the projection lives in a separate read-optimized store), the projection becomes eventually consistent with the log, which must be an explicit, accepted design point, not a surprise.
- Snapshotting. An order with a handful of events replays cheaply; an aggregate with thousands of events (unlikely for
Order, common for a long-lived aggregate elsewhere in a real system) needs periodic snapshots (a serialized current-state checkpoint plus “replay only events after sequence N”) to keep rebuild cost bounded — not implemented in this chapter’s example because Northwind’s orders don’t yet need it, but worth flagging as the pattern’s known scaling answer. - Event schema evolution. A field added to a new event type is safe; changing the meaning or shape of a past event type requires an upcasting layer (translating old event versions into the current in-memory shape during replay) since past events are never rewritten. Plan for this from the first event type onward — it’s cheaper to design an upcasting seam early than to retrofit one after the schema has already drifted.
- API versioning and backward compatibility. Unaffected at the REST/gRPC contract level — the projection preserves the same external shape used since Chapter 1, which is precisely why this migration didn’t require any consumer-facing change.
- Authentication and service-to-service trust. No new surface area here beyond what Chapters 4–6 already established — the event store is purely internal to
order-service. - Logging, metrics, tracing, audit trails. The event log is the audit trail for this service now — resist the temptation to also maintain a separate audit log; that would be duplicating the pattern’s core benefit into a second system that can drift from the first.
- Kubernetes deployment, health probes. No new infrastructure beyond
order-service’s existing Postgres instance (Chapter 3) — this chapter’s implementation deliberately avoids introducing a specialized event-store product, keeping the operational footprint unchanged. - Testing strategy. Test aggregate reconstruction (
Order.rebuild) as pure, fast unit tests over a list of events — no database needed for the core business-rule logic, a genuine testing advantage of this pattern once the event log and rebuild logic are separated cleanly, as shown above. - Migration strategy. Northwind migrated
order-servicealone, keeping its external contracts frozen, and back-filledorder_eventsfrom the existingorderstable’s audit columns where possible (an imperfect but pragmatic starting history) rather than attempting to reconstruct perfect historical events retroactively — a known, documented limitation of the migration, not a claim of complete historical fidelity.
8. Common Mistakes
- Applying event sourcing to every entity because one entity benefited from it. Migrating
inventory-service’s stock table to event sourcing “for consistency” adds real complexity with no corresponding auditability requirement. Fix: evaluate each aggregate independently against Section 3’s indicators; most of Northwind’s system correctly remains plain CRUD. - Querying the raw event log for read paths instead of a projection. Replaying every event on every
GET /orders/{id}call defeats read performance and turns a simple lookup into an expensive operation as history grows. Fix: always read through a maintained projection; reserve full replay for rebuilding that projection or for genuine “what happened historically” queries. - No optimistic-concurrency check on append. Allowing two concurrent commands to append conflicting events without detecting the race produces a corrupted, ambiguous history — worse than a lost update in a mutable table, because now the log itself contains an inconsistency. Fix: enforce a sequence-number check on every append, as shown in Section 5.
- Treating past events as mutable when the schema needs to change. Directly editing historical rows in
order_eventsto “fix” an old event’s shape corrupts the log’s fundamental guarantee — that it’s an immutable, trustworthy record. Fix: add an upcasting layer that translates old event shapes during replay; never rewrite history in place. - No snapshotting strategy planned before it’s needed. Discovering that a long-lived aggregate takes seconds to rebuild only after it has accumulated years of events, with no snapshot mechanism designed in, forces an urgent retrofit under production pressure. Fix: design the snapshot seam (even if unused initially, as in this chapter) as part of the initial event-sourcing implementation.
- Duplicating the event log as a separate audit log “to be safe.” Maintaining both the event-sourced log and a conventional audit trail means two systems that can drift and disagree about history. Fix: the event log is the single source of truth for history; don’t build a second one.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| Named compliance/audit requirement for full history of state transitions | Yes | Preserves complete history by construction, not as an afterthought | — |
| Domain has meaningful, named business events a domain expert recognizes | Yes | Event sourcing fits naturally when the domain is already event-shaped | — |
| Only need “who changed what, when” for occasional debugging | No | Full replay/projection machinery is overkill for a simple audit need | A conventional audit-log table alongside a mutable row |
| Entity has no real history requirement (e.g., a simple counter) | No | Adds replay and projection complexity for no corresponding benefit | Plain mutable row (Chapter 1 style) |
| Team has no prior experience and is under significant delivery pressure | No, not yet | Learning curve and schema-design risk are real; a mutable row is safer under pressure | Ship the mutable-row version first; revisit if the auditability need becomes concrete |
10. Hands-On Exercise
Extend it: implement a GET /orders/{id}/history endpoint that returns the full, human-readable sequence of events for an order — exactly the capability that would have resolved Section 1’s customer dispute in minutes instead of being unanswerable.
Simulate a failure: issue two concurrent cancel requests for the same order (from two separate threads or test clients) and confirm exactly one succeeds and the other receives a ConcurrentModificationException — verifying the optimistic-concurrency guard actually does what Section 5 claims.
Decision question, with justification required: Northwind’s finance team now wants the same point-in-time history capability for payment-service’s charge records, arguing “if it was good for orders, it’s good for payments.” Evaluate this request against Section 3’s indicators specifically for payment-service’s actual data and requirements (not by analogy to order-service) and justify your recommendation.
11. Key Takeaways
- Event sourcing stores state as an append-only sequence of meaningful domain events, with current state always derived by replay, never stored as the sole source of truth.
- It’s an internal persistence decision for one service at a time — exactly like database-per-service (Chapter 3), it doesn’t need to be, and usually shouldn’t be, applied uniformly across every service in a system.
- The pattern earns its complexity through a genuine, named requirement for auditable history or point-in-time queries — not general appeal or consistency with another service’s choice.
- Event sourcing (internal persistence) and event-driven architecture (Chapter 7, inter-service communication) are independent but complementary — a domain event can serve both roles at once, as this chapter’s
OrderCancelleddoes. - Always read through a maintained projection, never by replaying the full log on every query — replay is for rebuilding projections and answering genuine historical questions, not the default read path.
- Plan for event-schema evolution (upcasting) and snapshotting from the start, even if unused initially — retrofitting either under production pressure, after the log has already grown, is far more painful.
- A simpler audit-log table alongside a mutable row solves “who changed what, when” for most services without event sourcing’s replay and projection machinery — reserve the full pattern for entities whose history is genuinely, structurally part of the business problem.