Series overview
Part 22 of 2879% complete
2026-07-08•18 min read

Observability: logs, metrics, traces, and correlation IDs

Every chapter since Chapter 4 has said some version of “propagate the correlation ID” as a passing requirement. This chapter is where that discipline actually gets implemented, end to end, across the gateway, the BFFs, and every service in Northwind’s system.

1. Problem the Pattern Solves

A customer reports their checkout “took forever” — no error, no failure, just slow. The on-call engineer opens order-service’s logs and finds the request, but the log line says nothing about what order-service was waiting on: was it the call to inventory-service? The legacy ambassador? The saga orchestrator waiting on payment-service? Each of those services has its own log stream, in its own format, with no shared identifier connecting a single customer’s one slow checkout across all of them. Reconstructing the actual timeline means grep-ing five services’ logs by approximate timestamp and hoping nothing else happened at the same moment to confuse the picture — exactly the debugging cost this series predicted back in Chapter 7 when it first introduced asynchronous, multi-service flows.

A second, related gap: order-service’s dashboards show p50/p99 latency for the whole /orders endpoint, but not which part of that latency — the gateway, the BFF’s aggregation fan-out (Chapter 20), the saga’s synchronous first step, or a slow database query — actually dominates. Without that breakdown, “make checkout faster” has no clear starting point.

Forces in tension:

  • Completeness vs. cost and noise. Logging, tracing, and metering everything in maximum detail provides the most diagnostic power but costs real money (storage, ingestion) and can bury the signal that actually matters in noise nobody reads.
  • Consistency vs. per-service autonomy. Observability data is only useful for cross-service debugging if every service emits it in a compatible, correlatable format — which requires some centrally-agreed convention, in tension with each team’s freedom to instrument their own service however they like.
  • Traces vs. metrics vs. logs — different tools for different questions. Metrics answer “how much/how often, in aggregate” cheaply, at scale; traces answer “what exactly happened to this one request, across every service it touched,” at higher per-request cost; logs answer “what exactly did this one component say happened,” in unstructured or semi-structured detail. Conflating them, or trying to make one do all three jobs, degrades all three.
  • Instrumentation overhead vs. invisibility of production behavior. Every span, every metric, every structured log field has some runtime cost — negligible individually, potentially real in aggregate at high request volume — that must be weighed against the alternative of not knowing what’s actually happening in production.

2. Core Idea

Observability, in the sense this chapter uses it, rests on three complementary signal types plus one connecting thread:

  • Logs — discrete, timestamped records of specific events within one service, ideally structured (JSON, with consistent field names) rather than free-form text, so they can be queried and correlated programmatically.
  • Metrics — aggregated, numeric measurements over time (request rate, error rate, latency percentiles, queue depth) — cheap to store and query at scale, but only tell you that something is wrong in aggregate, not why for any specific request.
  • Traces — a record of one request’s full journey across every service it touched, broken into spans (one span per unit of work — an HTTP call, a database query, a Kafka publish) — the tool that actually answers “what happened to this specific slow checkout.”
  • Correlation ID — a single identifier generated at the point a request enters the system (the API gateway, per Chapter 4) and propagated through every subsequent hop — the thread that ties a request’s logs, its trace, and (where relevant) its metrics together into one coherent story.

trace-id: abc123

(generated here)

propagates trace-id

propagates trace-id

propagates trace-id

propagates trace-id

span: gateway routing, 5ms

span: aggregation fan-out, 180ms

span: reserveStock call, 2100ms <- the actual slow part

Client

API Gateway

mobile-bff

order-service

inventory-service

Saga: payment-service

OTel Collector

Trace backend

(e.g., Jaeger/Tempo)

trace-id: abc123

(generated here)

propagates trace-id

propagates trace-id

propagates trace-id

propagates trace-id

span: gateway routing, 5ms

span: aggregation fan-out, 180ms

span: reserveStock call, 2100ms <- the actual slow part

Client

API Gateway

mobile-bff

order-service

inventory-service

Saga: payment-service

OTel Collector

Trace backend

(e.g., Jaeger/Tempo)

Participants:

  • OpenTelemetry (OTel) — the instrumentation standard this chapter uses throughout, chosen specifically because it’s vendor-neutral: Northwind’s chosen trace backend can change later without re-instrumenting every service, unlike a vendor-proprietary SDK.
  • The OTel Collector — a separate process that receives telemetry from every service and forwards it to whichever backend(s) Northwind chooses (a trace store, a metrics store, a log aggregator) — decoupling instrumentation from backend choice.
  • Span context propagation — the mechanism (HTTP headers for REST/gRPC, message headers for Kafka) that carries the trace ID and span ID across every hop, including the gRPC hop from Chapter 6 and the Kafka hops from Chapters 7–8, each requiring its own explicit propagation wiring.

Commonly confused with:

  • APM (Application Performance Monitoring) as a single product. Commercial APM tools often bundle logs, metrics, and traces into one vendor’s platform — convenient, but this chapter deliberately uses the vendor-neutral OpenTelemetry standard for instrumentation, so Northwind’s data collection is decoupled from any specific backend’s product decisions.
  • Distributed tracing and correlation IDs as the same thing. A correlation ID is one field threaded through logs and requests; a distributed trace is a richer, structured graph of timed spans with parent-child relationships — a trace subsumes what a bare correlation ID provides, but a correlation ID alone (without full tracing infrastructure) is still valuable and much cheaper to implement, which is why this series has required it since Chapter 4 even before this chapter’s full tracing setup existed.
  • Health checks (Chapter 1 onward). A Kubernetes readiness/liveness probe answers “is this instance currently able to serve traffic” — a narrow, binary signal. Observability answers much richer questions about how the system is behaving across many requests and services — related, but not a substitute for each other.

3. When to Use It

Strong indicators:

  • More than one or two services in a request’s path (true for essentially every flow in this series since Chapter 4) — the point at which “read the one service’s logs” stops being sufficient to understand behavior.
  • A demonstrated debugging cost from lacking cross-service visibility — Section 1’s slow-checkout investigation, requiring manual log correlation across five services by timestamp, is exactly this evidence.
  • A performance question that requires knowing where time is spent across a multi-hop request, not just that the aggregate is slow.

Concrete use cases:

  • Any multi-service request path, which describes nearly every flow this series has built since Chapter 2 — order placement, the saga from Chapter 11, the aggregation from Chapter 20 all specifically benefit from end-to-end tracing.
  • Incident response of any kind: “why did this fail” and “why was this slow” are exactly the questions traces and correlated logs answer fastest, versus reconstructing a timeline from disconnected log streams under incident pressure.
  • Capacity planning and SLA reporting: metrics (aggregated latency percentiles, error rates) are the right tool for “how are we doing overall,” distinct from tracing’s per-request depth.
  • Compliance and audit trails: some regulated contexts require demonstrable request-level traceability across a distributed system — a coherent trace, not scattered logs, is often what satisfies this.

Prerequisites:

  • Agreement across teams on a shared correlation-ID/trace-context convention (this chapter standardizes on OpenTelemetry’s W3C Trace Context headers) — inconsistent conventions across services defeat the entire point.
  • Explicit propagation wiring at every non-HTTP boundary — Kafka message headers, gRPC metadata (Chapter 6) — since trace context doesn’t propagate automatically across those without deliberate instrumentation.
  • A backend capable of storing and querying traces and metrics at the retention and volume Northwind actually needs — and a sampling strategy (Section 7) once volume makes storing every single trace impractical.

4. When Not to Use It

  • A single-service system with no distributed request path. Full distributed tracing infrastructure for a system that never crosses a service boundary is solving a problem that doesn’t exist yet — good structured logging and basic metrics may be entirely sufficient.
  • Tracing every request at 100% sampling, indefinitely, at high volume, without a cost/value assessment. Storing a complete trace for every single request at Northwind’s eventual production scale could become a significant, unbounded cost — Section 7 addresses sampling directly as the answer, but it’s a decision to make deliberately, not default into by ignoring cost.
  • Using traces where a metric would answer the question more cheaply. “What’s our overall error rate this week” doesn’t need per-request trace inspection — a metric answers it directly and far more cheaply; reach for a trace specifically when the question is about one request’s specific path.
  • Overengineering signal: instrumenting every trivial internal method call as its own span, producing traces so deep and noisy that the actually-slow part is buried among hundreds of microsecond-scale spans nobody needed. Instrument at meaningful boundaries — HTTP calls, database queries, message publishes — not arbitrary internal function calls.

5. Implementation Example

OpenTelemetry auto-instrumentation, added to every service with minimal code changes — the deliberate design of OTel’s Java/Kotlin agent:

order-service/build.gradle.kts
dependencies {
implementation("io.opentelemetry.instrumentation:opentelemetry-spring-boot-starter:2.9.0-alpha")
implementation("io.micrometer:micrometer-tracing-bridge-otel")
implementation("io.opentelemetry:opentelemetry-exporter-otlp")
}
order-service/src/main/resources/application.yml
management:
tracing:
sampling:
probability: 1.0 # 100% in staging; see Section 7 for production sampling strategy
otlp:
tracing:
endpoint: http://otel-collector.observability.svc.cluster.local:4318/v1/traces

Auto-instrumentation gives order-service automatic spans for every incoming HTTP request, every outgoing RestClient call, and every JDBC query — no manual span creation needed for the common cases, which is most of what actually matters for diagnosing Section 1’s slow checkout.

Explicit correlation ID propagation at the gateway — the origin point, per Chapter 4’s original requirement, now actually implemented:

api-gateway/src/main/kotlin/in/o612/eng/northwind/gateway/TraceContextFilter.kt
package `in`.o612.eng.northwind.gateway
import org.springframework.cloud.gateway.filter.GlobalFilter
import org.springframework.core.Ordered
import org.springframework.stereotype.Component
import reactor.core.publisher.Mono
import java.util.UUID
@Component
class TraceContextFilter : GlobalFilter, Ordered {
override fun filter(exchange: org.springframework.web.server.ServerWebExchange, chain: org.springframework.cloud.gateway.filter.GatewayFilterChain): Mono<Void> {
// OTel's own instrumentation generates the actual W3C traceparent
// header automatically; this adds Northwind's own business-level
// correlation ID (useful in logs even without a full trace lookup).
val correlationId = exchange.request.headers.getFirst("X-Correlation-Id") ?: UUID.randomUUID().toString()
val mutatedRequest = exchange.request.mutate().header("X-Correlation-Id", correlationId).build()
return chain.filter(exchange.mutate().request(mutatedRequest).build())
}
override fun getOrder() = Ordered.HIGHEST_PRECEDENCE
}

Propagating trace context across the Kafka boundary — the specific gap OTel’s HTTP auto-instrumentation doesn’t cover automatically, requiring explicit wiring, exactly as Chapters 7 and 19 flagged but deferred:

order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/outbox/OutboxRelay.kt (revised)
package `in`.o612.eng.northwind.order.internal.outbox
import io.opentelemetry.api.trace.Span
import io.opentelemetry.context.Context
import io.opentelemetry.context.propagation.TextMapSetter
import org.springframework.kafka.core.KafkaTemplate
import org.springframework.kafka.support.KafkaHeaders
import org.springframework.messaging.support.MessageBuilder
class OutboxRelay(private val kafkaTemplate: KafkaTemplate<String, String>, private val otelPropagator: io.opentelemetry.context.propagation.TextMapPropagator) {
fun relay(aggregateId: String, topic: String, payload: String) {
val messageBuilder = MessageBuilder.withPayload(payload)
.setHeader(KafkaHeaders.TOPIC, topic)
.setHeader(KafkaHeaders.KEY, aggregateId)
// Inject the current trace context into Kafka message headers —
// without this, every event-driven flow (Chapters 7-12) becomes
// an untraceable gap: the trace simply ends at the publish and a
// new, disconnected one starts at each consumer.
otelPropagator.inject(Context.current(), messageBuilder) { builder, key, value ->
builder?.setHeader(key, value.toByteArray())
}
kafkaTemplate.send(messageBuilder.build())
}
}
inventory-service/src/main/kotlin/in/o612/eng/northwind/inventory/internal/OrderPlacedListener.kt (revised)
package `in`.o612.eng.northwind.inventory.internal
import io.opentelemetry.context.Context
import io.opentelemetry.context.propagation.TextMapGetter
import org.springframework.kafka.annotation.KafkaListener
import org.springframework.kafka.support.KafkaHeaders
import org.springframework.messaging.handler.annotation.Headers
import org.springframework.stereotype.Component
@Component
class OrderPlacedListener(private val otelPropagator: io.opentelemetry.context.propagation.TextMapPropagator, private val inventoryService: InventoryService) {
@KafkaListener(topics = ["orders.events"], groupId = "inventory-service")
fun onOrderPlaced(event: `in`.o612.eng.northwind.order.api.OrderPlaced, @Headers headers: Map<String, Any>) {
// Extract the propagated trace context so this consumer's spans
// attach to the SAME trace the original checkout request started —
// the specific fix that makes Chapter 7's event-driven flows
// traceable end to end instead of appearing as disconnected traces.
val extractedContext = otelPropagator.extract(Context.current(), headers) { h, key ->
(h[key] as? ByteArray)?.toString(Charsets.UTF_8)
}
extractedContext.makeCurrent().use {
inventoryService.reserveStock(event.orderId, event.items.toReservationRequests())
}
}
}

A custom span for the specific slow operation from Section 1, giving the trace explicit visibility into exactly what Section 1’s investigation needed:

order-service/src/main/kotlin/in/o612/eng/northwind/order/internal/InventoryClient.kt (span annotation added)
package `in`.o612.eng.northwind.order.internal
import io.opentelemetry.instrumentation.annotations.WithSpan
import io.opentelemetry.instrumentation.annotations.SpanAttribute
class InventoryClient(/* ... */) {
@WithSpan("inventory.reserveStock")
fun reserveStock(@SpanAttribute("order.id") orderId: java.util.UUID, items: List<ReservationItemDto>) {
// existing implementation from Chapters 2, 16 — now producing a
// named, attributed span visible in the trace as its own timed segment
}
}

6. Step-by-Step Flow

Revisiting Section 1’s slow-checkout investigation, now with the tooling this chapter builds:

inventory-service (via Kafka)order-servicemobile-bffAPI GatewayClientinventory-service (via Kafka)order-servicemobile-bffAPI GatewayClientOne trace, one ID, spans every hop —the on-call engineer sees exactly whichspan took 2100ms, in one queryPOST /orders (trace started, trace-id=abc123)forward (trace-id propagated)aggregate call (trace-id propagated)span: reserveStock, 2100ms <- visible immediately in the tracepublish OrderPlaced (trace context injected into Kafka headers)span: reserveStock (consumer side), same trace-id=abc123
inventory-service (via Kafka)order-servicemobile-bffAPI GatewayClientinventory-service (via Kafka)order-servicemobile-bffAPI GatewayClientOne trace, one ID, spans every hop —the on-call engineer sees exactly whichspan took 2100ms, in one queryPOST /orders (trace started, trace-id=abc123)forward (trace-id propagated)aggregate call (trace-id propagated)span: reserveStock, 2100ms <- visible immediately in the tracepublish OrderPlaced (trace context injected into Kafka headers)span: reserveStock (consumer side), same trace-id=abc123
  1. Client action. The same checkout request from Section 1, now instrumented end to end.
  2. API request. The gateway generates (or propagates) both the OTel trace context and Northwind’s own correlation ID, attaching both to every downstream hop.
  3. Service behavior. Each service’s auto-instrumentation creates spans for its own HTTP handling and outbound calls, with no manual code beyond the explicit Kafka propagation shown above.
  4. Database interaction. JDBC auto-instrumentation captures each query as its own span — directly answering “was this a slow database query” without guessing.
  5. Inter-service communication. The Kafka publish and consume, explicitly instrumented (Section 5), keep the trace continuous across the asynchronous boundary that would otherwise have silently broken it.
  6. Error or failure handling. A failed span (an exception, a non-2xx response) is marked as an error in the trace, visually distinct from a merely slow one — directly distinguishing “this failed” from “this was just slow,” a distinction raw logs often blur.
  7. Observability signals. The trace backend now shows, for this exact request, that inventory.reserveStock’s span took 2100ms while every other span took under 50ms — answering Section 1’s original question in one query instead of a multi-service log archaeology exercise.
  8. Final response/outcome. The on-call engineer identifies the actual bottleneck (a slow query inside reserveStock, visible as a nested database span) in minutes, not the hours Section 1’s original manual correlation would have taken.

7. Production Concerns

  • Sampling strategy. 100% trace sampling (Section 5’s staging configuration) becomes cost-prohibitive at production volume — a common approach is head-based sampling (sample a percentage of requests at the entry point, propagate that decision through the whole trace) or tail-based sampling (capture everything temporarily, retain only traces that were slow or errored) — Northwind should choose deliberately based on actual trace-backend cost, not default to either extreme.
  • Data consistency. Unaffected — observability is a read-only, side-effect-free concern layered onto existing request flows.
  • API versioning. Not directly relevant, though span names and metric names should themselves be treated with some naming-convention discipline (this series’ consistent service.operation pattern, e.g., inventory.reserveStock) so dashboards and alerts don’t need updating every time an internal refactor happens.
  • Authentication and service-to-service trust. The OTel Collector and trace/metrics backends need their own access control — trace data can contain sensitive request details (the @SpanAttribute("order.id") above is fine; a customer’s email address as a span attribute would not be, and needs the same data-sensitivity review as any log field).
  • Logging, metrics, tracing, correlation IDs (this chapter’s own subject): structure logs as JSON with the trace ID as a standard field in every log line — this is what lets a log query and a trace lookup for the same request cross-reference each other directly, the actual mechanism that ties logs and traces into one coherent investigation tool.
  • Kubernetes deployment, health probes, autoscaling. The OTel Collector itself needs to be deployed with its own availability and scaling plan — a collector that falls behind or crashes silently drops telemetry, which (unlike a crashed application) often goes unnoticed until someone needs a trace that was never captured.
  • Testing strategy. Verify trace propagation explicitly in integration tests — assert that a request’s trace ID appears consistently across every service’s logs in a test scenario, rather than assuming instrumentation “just works” once configured; Kafka header propagation (Section 5) specifically deserves its own dedicated test, since it’s the easiest link in the chain to silently break.
  • Migration strategy. Instrument the highest-value flows first — Northwind started with the checkout path given Section 1’s actual incident — and expand coverage incrementally, service by service, rather than attempting a simultaneous cluster-wide instrumentation rollout.

8. Common Mistakes

  1. No trace context propagation across Kafka. Assuming OTel’s HTTP auto-instrumentation covers the entire request path, when every Kafka publish/consume boundary silently starts a new, disconnected trace unless explicitly wired (Section 5). Fix: explicitly inject and extract trace context at every message-broker boundary, as shown.
  2. 100% sampling at production scale with no cost review. Enabling full tracing in production without considering storage and ingestion cost at real traffic volume can produce an unpleasant budget surprise. Fix: choose and implement a deliberate sampling strategy (Section 7) before production rollout, not after the first large bill.
  3. Putting sensitive data in span attributes or log fields without review. Adding a customer’s email address, payment details, or other sensitive data as a trace attribute “for debugging convenience” creates a compliance and security exposure inside observability tooling that may have different access controls than the application itself. Fix: review span attributes and log fields for sensitivity with the same rigor as any other data-handling decision.
  4. Instrumenting every trivial method call as its own span. Over-instrumentation produces traces so deep and noisy that the genuinely slow span is buried among hundreds of microsecond-scale ones. Fix: instrument at meaningful boundaries (HTTP, database, messaging) as auto-instrumentation already does well; add manual spans (Section 5’s @WithSpan) only for specific, meaningful operations worth seeing separately.
  5. Inconsistent correlation ID or trace-context conventions across teams. If one team’s service generates its own ad hoc “request ID” instead of respecting the propagated OTel trace context, that service’s logs can’t be correlated with the rest of the trace. Fix: standardize on one convention (W3C Trace Context, as this chapter does) platform-wide, enforced through shared library defaults rather than left to each team’s independent choice.
  6. Treating an observability rollout as “done” once instrumentation code ships. Shipping the instrumentation without verifying, end to end, that a real request’s trace actually appears correctly connected across every hop it should — as this chapter’s Section 6 walkthrough does — risks discovering the gaps only during a real incident, exactly when they’re most costly to find. Fix: verify end-to-end trace continuity as an explicit acceptance test for the observability rollout itself, not an assumption.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Requests span multiple services, debugging requires manual log correlation across themYesDistributed tracing directly answers “what happened to this request, everywhere”—
Need aggregate answers (“what’s our error rate this week”)Metrics, not full tracing for this questionCheaper, purpose-built for aggregate answersMetrics dashboards
Single-service system, no distributed request pathNo (full tracing)Nothing to correlate across yetStructured logging + basic metrics
High request volume, full trace retention is cost-prohibitiveYes, with samplingA deliberate sampling strategy preserves diagnostic value at sustainable costHead-based or tail-based sampling, chosen deliberately
An asynchronous (Kafka) hop is part of the flow being investigatedYes, with explicit context propagationWithout it, the trace silently breaks at every message-broker boundary—

10. Hands-On Exercise

Extend it: add explicit trace-context propagation to the gRPC path from Chapter 6 (warehouse-bff’s call to inventory-service’s StockLookupService), verifying the resulting trace connects correctly across that transport, exactly as Section 5 did for Kafka.

Simulate a failure: deliberately skip Section 5’s Kafka header propagation for one event type, and observe what the resulting trace looks like for a flow crossing that specific boundary — confirming firsthand what “the trace silently breaks” actually looks like in a real trace backend, so you recognize it during a future investigation.

Decision question, with justification required: Northwind’s trace backend costs are rising as request volume grows, and the platform team is deciding between head-based sampling (a fixed 10% of requests, decided at the gateway) and tail-based sampling (capture everything temporarily, retain only traces that errored or exceeded a latency threshold). Given that Section 1’s original incident was a slow, not failed request, which sampling strategy would have reliably captured it, and what does that imply about which approach better serves Northwind’s actual debugging needs?

11. Key Takeaways

  • Logs, metrics, and traces answer different questions — logs, “what did this component say”; metrics, “how much/how often, in aggregate”; traces, “what happened to this one request, everywhere it went” — and conflating them, or using one where another fits better, degrades all three.
  • A correlation ID and full distributed tracing are related but distinct investments — this series required correlation-ID propagation since Chapter 4, well before this chapter’s full OpenTelemetry tracing infrastructure, because even a bare correlation ID provides real value at much lower implementation cost.
  • Explicit trace-context propagation is required at every non-HTTP boundary — Kafka message headers, gRPC metadata — since auto-instrumentation’s HTTP coverage doesn’t extend there without deliberate wiring, and skipping it silently breaks the trace at exactly the asynchronous boundaries this series has built since Chapter 7.
  • Choose a deliberate sampling strategy before production scale makes 100% trace retention cost-prohibitive — and choose it based on what kind of incidents (slow vs. failed, as Section 10 asks) it actually needs to reliably capture.
  • Review every span attribute and log field for sensitive data before it ships — observability tooling can become an unreviewed side channel for exactly the data your application’s own access controls were built to protect.
  • Instrument at meaningful boundaries (HTTP, database, messaging), not every internal method call — over-instrumentation buries the genuinely useful signal in noise.
  • Verify end-to-end trace continuity as an explicit acceptance test for any observability rollout — assuming instrumentation “just works” once configured risks discovering real gaps only during an actual incident.
Spring BootKotlinMicroservicesObservability

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind