Series overview
Part 25 of 2889% complete
2026-07-13•18 min read

Kubernetes-native microservice design

Every chapter since Chapter 1 has used a Kubernetes primitive in passing — a readiness probe here, a Service object there, an HorizontalPodAutoscaler in Chapter 18. This chapter is the one that steps back and asks: what does it actually mean for a Spring Boot service to be designed for Kubernetes, rather than merely deployed onto it?

1. Problem the Pattern Solves

A new engineer joins Northwind’s platform team and, reviewing payment-service’s deployment configuration, finds it works — mostly. Pods restart cleanly on a rolling update. But digging further: the readiness probe checks only /actuator/health, which reports healthy even when payment-service’s connection pool to its database is fully exhausted (Spring Boot’s default health indicator doesn’t check pool saturation) — meaning Kubernetes keeps routing traffic to a pod that’s actually unable to serve requests promptly. Startup takes eleven seconds, but the readiness probe’s initialDelaySeconds is set to three, so during every rolling deploy, a short window exists where Kubernetes considers a still-initializing pod ready and sends it traffic that fails. And the Deployment’s terminationGracePeriodSeconds is the Kubernetes default of 30 seconds, but nothing in payment-service’s code actually listens for SIGTERM to stop accepting new requests and finish in-flight ones gracefully — a pod being terminated during a deploy can drop requests mid-flight.

None of these are exotic problems. They’re the ordinary gap between “runs on Kubernetes” and “designed with Kubernetes’ actual operational model in mind” — a gap most services fall into by default, because the failure modes only appear under real deployment churn and real load, not in a quick local test.

Forces in tension:

  • Convenience of defaults vs. correctness under Kubernetes’ actual operational model. Spring Boot’s and Kubernetes’ defaults are reasonable starting points, but neither is tuned to the other’s specific expectations by default — the two systems’ assumptions must be deliberately reconciled, not assumed compatible.
  • Graceful shutdown vs. deployment speed. Waiting for in-flight requests to complete before terminating a pod is correct but takes time — a deploy that impatiently kills pods for speed risks dropped requests; one that waits too generously slows every rollout.
  • Accurate health signaling vs. probe overhead. A health check thorough enough to catch real problems (database pool saturation, downstream dependency failures) costs more to compute than a bare “is the process running” check — but a check that’s too shallow gives Kubernetes false confidence in a pod that can’t actually serve traffic.
  • Resource requests/limits accuracy vs. operational safety margin. Under-requesting resources risks a pod being starved or OOM-killed under real load; over-requesting wastes cluster capacity and money — getting this right requires real production data, not a guess made once at initial deployment.

2. Core Idea

A Kubernetes-native service is one whose lifecycle hooks, health signaling, resource declarations, and configuration are deliberately designed around Kubernetes’ actual operational model — not just a service that happens to run inside a container Kubernetes schedules. This chapter consolidates the specific mechanisms:

Pod lifecycle, correctly wired

pod removed from Service

endpoints BEFORE this fires

Startup probe:

waits for full readiness

before liveness/readiness even begin

Liveness probe:

is the process healthy,

should it be restarted

Readiness probe:

can it serve traffic RIGHT NOW

(checks real dependencies)

SIGTERM handling:

stop accepting new requests,

finish in-flight ones,

THEN exit

Pod lifecycle, correctly wired

pod removed from Service

endpoints BEFORE this fires

Startup probe:

waits for full readiness

before liveness/readiness even begin

Liveness probe:

is the process healthy,

should it be restarted

Readiness probe:

can it serve traffic RIGHT NOW

(checks real dependencies)

SIGTERM handling:

stop accepting new requests,

finish in-flight ones,

THEN exit

Participants:

  • Startup, liveness, and readiness probes — three distinct signals, each answering a different question, all needed together and often conflated into one under-specified health check.
  • Graceful shutdown handling — explicit application-level response to SIGTERM, coordinated with Kubernetes’ own pod-termination sequence (removing the pod from a Service’s endpoints before sending SIGTERM, not simultaneously).
  • Resource requests and limits — Kubernetes’ scheduling and QoS mechanism, requiring real data to set correctly rather than defaults or guesses.
  • ConfigMaps and Secrets — the Kubernetes-native mechanisms for configuration and credentials, distinct from (and sometimes complementary to) the Spring Cloud Config server from Chapter 5.

Commonly confused with:

  • “Runs in a container” or “runs on Kubernetes.” Every service in this series has run in containers since Chapter 1 — that alone says nothing about whether its lifecycle hooks, health checks, and resource declarations are actually correct for Kubernetes’ operational model, as payment-service’s gaps in Section 1 demonstrate.
  • Cloud-native in the broader marketing sense. “Cloud-native” is sometimes used to mean “uses microservices, containers, and CI/CD” generally — this chapter’s narrower, concrete meaning is specifically about correct integration with Kubernetes’ pod lifecycle and scheduling model, a subset of that broader term with precise, testable criteria.
  • Service mesh adoption (Chapter 21). A mesh adds a traffic-management and security layer on top of Kubernetes; Kubernetes-native design (this chapter) is about correctly using Kubernetes’ own primitives underneath, independent of whether a mesh is present — a mesh doesn’t fix an incorrect readiness probe or a missing graceful-shutdown hook.

3. When to Use It

Strong indicators:

  • Any service deployed to Kubernetes, which describes every service in this series since Chapter 1 — this isn’t an optional add-on pattern for special cases, it’s the baseline correctness bar for running well on the platform this series has assumed throughout.
  • Observed or plausible symptoms like Section 1’s: dropped requests during deploys, pods marked ready before they can actually serve traffic, or pods restarted (or not restarted) at the wrong times relative to their actual health.
  • A team that has only ever used Kubernetes’ defaults without deliberately reconciling them against Spring Boot’s own startup, shutdown, and health-reporting behavior.

Concrete use cases:

  • Any production Kubernetes deployment, without exception — this is foundational hygiene, not a specialized pattern reserved for particular domains, unlike most patterns in this series.
  • High-deploy-frequency platforms: a team deploying multiple times a day feels the cost of incorrect graceful shutdown and readiness handling far more acutely than one deploying monthly, since every deploy is a chance to drop requests.
  • Cost-sensitive platforms: correctly tuned resource requests/limits (Section 5) directly affect cluster cost and bin-packing efficiency at any meaningful scale.
  • Regulated or high-availability services: a payment or healthcare system where a dropped request during a routine deploy has real consequences needs this discipline especially rigorously.

Prerequisites:

  • Actual production load data (or a realistic load test) to set resource requests/limits meaningfully — guessing these values without data produces exactly the under- or over-provisioning risk Section 1 describes.
  • A health-check design that reflects real dependency health (database connectivity, critical downstream calls) rather than just process liveness — Spring Boot Actuator’s health indicators need to be configured deliberately, not accepted at their defaults.
  • Explicit SIGTERM handling wired into the application, not assumed to happen automatically — Spring Boot’s graceful shutdown support exists but must be enabled and tuned.

4. When Not to Use It

There’s no legitimate “not needed” case for a production Kubernetes deployment — every service benefits from correct lifecycle and health-signaling design. The judgment calls are about depth and precision, not whether to bother at all:

  • A short-lived experimental service or spike may reasonably defer fine-tuned resource limits and elaborate health checks until it’s clear the service will persist — spending a day tuning terminationGracePeriodSeconds for a throwaway prototype is effort better spent elsewhere.
  • Overengineering signal: building an elaborate, custom health-check endpoint checking dozens of downstream dependencies’ full round-trip latency on every readiness probe call (which Kubernetes calls frequently) can itself become a performance and cascading-failure risk — a readiness probe should be fast and cheap, checking the dependencies that genuinely determine “can this pod serve traffic right now,” not exhaustively probing everything the service ever calls.
  • Risk of over-precise resource limits set too early. Setting very tight CPU/memory limits based on a single load test, before understanding real production traffic variance, risks pods being OOM-killed or CPU-throttled under legitimate but unanticipated load patterns — start with reasonable headroom and tighten based on sustained production observation, not a single benchmark.

5. Implementation Example

Three distinct probes, each answering its own specific question — the direct fix for Section 1’s conflated single health check:

k8s/payment-service-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-service
spec:
template:
spec:
terminationGracePeriodSeconds: 45 # generous enough for in-flight requests to finish — see graceful shutdown below
containers:
- name: payment-service
image: northwind/payment-service:3.2.0
resources:
requests: { cpu: "500m", memory: "768Mi" } # set from real production observation, not a guess
limits: { cpu: "1000m", memory: "1024Mi" }
startupProbe:
httpGet: { path: /actuator/health/readiness, port: 8080 }
failureThreshold: 30 # allow up to 30 * 2s = 60s for slow startup (JIT warmup, connection pool init)
periodSeconds: 2
livenessProbe:
httpGet: { path: /actuator/health/liveness, port: 8080 }
periodSeconds: 10
failureThreshold: 3 # restart only after sustained failure, not one transient blip
readinessProbe:
httpGet: { path: /actuator/health/readiness, port: 8080 }
periodSeconds: 5
failureThreshold: 2 # remove from traffic quickly if it can't actually serve requests

The startupProbe is what fixes Section 1’s “eleven seconds to start, three-second initialDelaySeconds” mismatch — liveness and readiness probes don’t even begin until the startup probe succeeds, so a genuinely slow-starting pod is never prematurely marked ready or killed for failing a liveness check before it’s finished initializing.

A readiness check that reflects real, actionable dependency health — not just “is the process alive,” per Section 1’s database-pool-saturation gap:

payment-service/src/main/kotlin/in/o612/eng/northwind/payment/internal/DatabasePoolHealthIndicator.kt
package `in`.o612.eng.northwind.payment.internal
import com.zaxxer.hikari.HikariDataSource
import org.springframework.boot.actuate.health.Health
import org.springframework.boot.actuate.health.HealthIndicator
import org.springframework.stereotype.Component
@Component
class DatabasePoolHealthIndicator(private val dataSource: HikariDataSource) : HealthIndicator {
override fun health(): Health {
val pool = dataSource.hikariPoolMXBean
val utilizationPercent = (pool.activeConnections.toDouble() / dataSource.maximumPoolSize) * 100
// Above 90% pool utilization, this pod genuinely cannot serve new
// requests promptly — readiness should reflect that reality, not
// report healthy just because the JVM process is technically alive.
return if (utilizationPercent > 90) {
Health.down().withDetail("poolUtilizationPercent", utilizationPercent).build()
} else {
Health.up().withDetail("poolUtilizationPercent", utilizationPercent).build()
}
}
}
payment-service/src/main/resources/application.yml
management:
endpoint:
health:
group:
readiness:
include: readinessState, db, databasePool # includes the custom indicator above
liveness:
include: livenessState # deliberately minimal — liveness should only catch "the process is truly stuck"

Liveness and readiness are configured to check different things, deliberately: liveness stays minimal (only Spring’s own internal liveness state — is the application context still functioning at all) because a liveness failure triggers a pod restart, an expensive and potentially disruptive action that should only happen for genuinely unrecoverable states; readiness includes the database pool check because a readiness failure only removes the pod from traffic temporarily, a much cheaper and more reversible action appropriate for a transient condition like pool saturation.

Graceful shutdown, actually wired to respond to SIGTERM correctly:

payment-service/src/main/resources/application.yml (continued)
server:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 40s # slightly less than terminationGracePeriodSeconds above, leaving margin

With server.shutdown: graceful enabled, Spring Boot stops accepting new requests on SIGTERM and waits (up to the configured timeout) for in-flight requests to complete before the JVM actually exits — combined with Kubernetes’ own behavior of removing a terminating pod from its Service’s endpoints before sending SIGTERM (the default, correct sequencing), in-flight requests finish and new requests stop arriving, with no dropped connections during a routine rolling deploy.

Resource requests set from real data, not a guess — querying Prometheus (Chapter 22’s infrastructure) for payment-service’s actual observed usage before setting limits:

Query used to inform the resource values in the Deployment above
quantile_over_time(0.95, container_memory_working_set_bytes{pod=~"payment-service-.*"}[30d])
quantile_over_time(0.95, rate(container_cpu_usage_seconds_total{pod=~"payment-service-.*"}[5m])[30d:])

Setting requests near the observed p95 and limits with meaningful headroom above it (not an arbitrary multiplier) reflects actual behavior rather than an initial guess made before the service had any real traffic history — exactly the kind of decision this series has insisted be evidence-driven throughout.

6. Step-by-Step Flow

Tracing a rolling deploy of payment-service, correctly configured, versus Section 1’s original gaps:

Service endpointspayment-service (new pod)payment-service (old pod)KubernetesService endpointspayment-service (new pod)payment-service (old pod)Kubernetes~11s: JVM starts, connection pool initializesZero dropped requests throughout —new pod never received traffic before truly ready;old pod never killed mid-requestcreate new podstartupProbe checks readiness/health repeatedlystartupProbe succeedsbegin liveness + readiness probingreadinessProbe succeeds (pool healthy)add new pod to endpointsremove old pod from endpoints (BEFORE termination)send SIGTERMstop accepting new requests, finish in-flight ones (up to 40s)exit cleanly
Service endpointspayment-service (new pod)payment-service (old pod)KubernetesService endpointspayment-service (new pod)payment-service (old pod)Kubernetes~11s: JVM starts, connection pool initializesZero dropped requests throughout —new pod never received traffic before truly ready;old pod never killed mid-requestcreate new podstartupProbe checks readiness/health repeatedlystartupProbe succeedsbegin liveness + readiness probingreadinessProbe succeeds (pool healthy)add new pod to endpointsremove old pod from endpoints (BEFORE termination)send SIGTERMstop accepting new requests, finish in-flight ones (up to 40s)exit cleanly
  1. Client action. A routine deploy of a payment-service patch release, no different in kind from any deploy this series has performed since Chapter 1.
  2. API request equivalent. Kubernetes creates the new pod and begins the startup probe cycle — no traffic reaches it yet.
  3. Service behavior. The new pod’s connection pool initializes fully before the startup probe succeeds, at which point liveness and readiness probing begin.
  4. Database interaction. The custom DatabasePoolHealthIndicator genuinely reflects the pool’s real state — readiness only succeeds once the pool is actually usable, not merely once the process has started.
  5. Inter-service communication. Unaffected — this pattern operates at the pod-lifecycle level, below any service-to-service call.
  6. Error or failure handling. If the new pod’s readiness check never succeeds (a genuine startup failure), Kubernetes never routes traffic to it and the rollout can be halted or rolled back — the old pods keep serving traffic the entire time, since Kubernetes only removes an old pod once its replacement is confirmed ready.
  7. Observability signals. Track probe failure counts and pod restart counts as explicit metrics (feeding directly into the resource-tuning feedback loop from Section 5) — a rising restart count on a specific service is a leading indicator worth investigating before it becomes a customer-visible incident.
  8. Final response/outcome. The rolling deploy completes with zero dropped requests and zero prematurely-routed traffic — the specific, concrete fix for every gap Section 1 identified in payment-service’s original configuration.

7. Production Concerns

  • Timeouts, retries, idempotency. terminationGracePeriodSeconds and Spring’s timeout-per-shutdown-phase must be coordinated (the Spring value should be somewhat less than the Kubernetes value, as Section 5 configures) — if Kubernetes force-kills the pod (SIGKILL) before Spring’s own graceful shutdown completes, in-flight requests are still dropped despite the graceful-shutdown configuration existing.
  • Data consistency. A request mid-database-transaction when SIGTERM arrives should either complete (within the graceful shutdown window) or roll back cleanly — verify this explicitly for any long-running operation, since Kubernetes’ termination sequence doesn’t know or care about application-level transaction boundaries.
  • API versioning. Not directly relevant, though a startup probe’s failure threshold should be generous enough to tolerate a legitimately slower-starting new version (a larger JIT warmup after a major dependency upgrade, for instance) without falsely flagging it as a failed deploy.
  • Authentication and service-to-service trust. Secrets should be mounted via Kubernetes Secret objects (ideally backed by an external secrets manager, as Chapter 5 discussed) rather than baked into container images or passed as plain environment variables in a ConfigMap — a Secret at minimum gets base64 encoding and Kubernetes’ own RBAC-scoped access control, neither of which a ConfigMap provides.
  • Logging, metrics, tracing, correlation IDs. Ensure logs are flushed and any final trace spans are exported before the process exits during graceful shutdown — an application that buffers logs or traces in memory and exits before flushing them loses exactly the diagnostic data most needed if the shutdown itself was triggered by a problem.
  • Kubernetes deployment, health probes, autoscaling. This entire chapter is this concern, consolidated — the specific combination of startup/liveness/readiness probes, graceful shutdown, and resource requests/limits, tuned together rather than independently, is what this pattern actually delivers.
  • Testing strategy. Test graceful shutdown explicitly — send a SIGTERM to a running instance mid-load-test and verify in-flight requests complete successfully while new ones are rejected or queued appropriately; this is a specific, valuable test most teams skip because it doesn’t fit neatly into a typical unit or integration test.
  • Migration strategy. Audit every existing service against this chapter’s checklist (three distinct probes, real dependency-aware readiness, wired graceful shutdown, data-driven resource limits) incrementally, prioritizing services with the highest deploy frequency or the most business-critical availability requirements first — exactly the evidence-driven prioritization this series has applied to every pattern.

8. Common Mistakes

  1. Using one health endpoint for both liveness and readiness. Pointing both probes at the same check conflates two different questions and two different remediation actions (restart vs. remove from traffic) — a transient dependency issue that should only affect readiness can trigger unnecessary pod restarts if it’s also wired to liveness. Fix: configure separate health groups, as Section 5 does, with liveness minimal and readiness dependency-aware.
  2. A readiness check that doesn’t reflect real serviceability. Checking only “is the HTTP server running” (Spring Boot’s bare default) misses exactly the database-pool-saturation scenario from Section 1 — the pod reports healthy while genuinely unable to serve requests promptly. Fix: include real, actionable dependency checks (connection pool utilization, critical downstream reachability) in the readiness group specifically.
  3. No startupProbe for a service with meaningfully slow initialization. Relying only on initialDelaySeconds on the liveness/readiness probes, set too short for the service’s actual startup time, causes false failures during every deploy of a slow-starting service. Fix: use a dedicated startupProbe with a generous failure threshold, as Section 5 configures, so liveness/readiness checking doesn’t even begin until startup genuinely completes.
  4. No graceful shutdown handling, relying on Kubernetes’ default SIGKILL after the grace period. Without server.shutdown: graceful enabled and correctly timed, in-flight requests are simply dropped when the pod terminates, regardless of how generous terminationGracePeriodSeconds is set. Fix: explicitly enable and tune Spring Boot’s graceful shutdown, coordinated with the Kubernetes-level grace period.
  5. Resource requests/limits set once, from a guess, and never revisited. Values chosen before any real production traffic existed often turn out significantly wrong once actual usage patterns emerge — either wasting cluster capacity or risking throttling/OOM kills under real load. Fix: revisit resource requests/limits periodically using real observed data (Section 5’s Prometheus query), treating them as a tuned, evolving value rather than a one-time setting.
  6. Overloading the readiness probe with expensive, exhaustive dependency checks. A readiness endpoint that makes full round-trip calls to every downstream service on every probe invocation (which Kubernetes calls frequently, every few seconds per pod) adds real, avoidable load and can itself become a cascading-failure vector under stress. Fix: keep readiness checks fast and focused on the specific dependencies that genuinely determine immediate serviceability, not an exhaustive health survey of the whole system.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Dropped requests or premature traffic routing during rolling deploysYesCorrect probe and graceful-shutdown configuration directly addresses this—
Pods restarted for transient issues that should only affect readinessYesSeparating liveness from readiness prevents unnecessary, disruptive restarts—
Resource requests/limits set once, before any real traffic history existedYes, revisit with dataGuessed values are frequently wrong once real usage patterns are known—
Short-lived experimental or throwaway serviceLighter touch is fineFull tuning effort isn’t justified for something that won’t persistBasic health check; defer full tuning until the service proves durable
Readiness check exhaustively probing every downstream dependency on every callNo, that’s overcorrectionAdds load and cascading-failure risk; keep readiness fast and focusedCheck only the dependencies that determine immediate serviceability

10. Hands-On Exercise

Extend it: apply this chapter’s full checklist (three distinct probes, dependency-aware readiness, graceful shutdown, data-driven resource limits) to inventory-service, using its own real Prometheus data (from Chapter 22’s instrumentation) rather than copying payment-service’s specific values verbatim — justify any differences based on inventory-service’s actual observed behavior.

Simulate a failure: run a load test against payment-service while triggering a rolling deploy mid-test, and measure the actual number of dropped or failed requests before and after applying this chapter’s graceful-shutdown and probe configuration — quantifying the concrete improvement rather than assuming it.

Decision question, with justification required: Northwind’s platform team is debating whether liveness probe failures should trigger an automatic pod restart immediately (failureThreshold: 1) or only after several consecutive failures (failureThreshold: 3, as Section 5 configures). What trade-off does this specific number represent, and what evidence (from Chapter 22’s observability data) would help decide it correctly for a specific service rather than guessing?

11. Key Takeaways

  • Being Kubernetes-native means deliberately designing a service’s lifecycle hooks, health signaling, and resource declarations around Kubernetes’ actual operational model — not merely running the service inside a container Kubernetes happens to schedule.
  • Liveness and readiness answer genuinely different questions with genuinely different remediation actions (restart vs. remove from traffic) — conflating them into one health check, as many services do by default, causes both unnecessary restarts and undetected unserviceable pods.
  • A startupProbe is what lets liveness and readiness checking correctly wait for a service’s actual, sometimes-slow initialization, instead of forcing a compromise between initialDelaySeconds values that are either too short (false failures) or too long (slow failure detection).
  • Graceful shutdown requires explicit application-level coordination with Kubernetes’ termination sequence — enabling it and tuning its timeout relative to terminationGracePeriodSeconds is what actually prevents dropped requests during routine deploys.
  • Resource requests and limits should be set from real, observed production data, not an initial guess — and revisited periodically as actual usage patterns become clear, since a value chosen before any real traffic existed is often wrong.
  • Keep readiness checks fast and focused on dependencies that genuinely determine immediate serviceability — an exhaustive, expensive readiness check can itself become a performance and cascading-failure risk under Kubernetes’ frequent probing.
  • This chapter isn’t a specialized pattern for particular domains — it’s the baseline correctness bar every service in this series (and any production Kubernetes deployment) should meet, regardless of which other patterns from this book it also uses.
Spring BootKotlinMicroservicesKubernetes

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind