Diagnosing common Spring Boot bottlenecks
By the end of this chapter you have a catalog of the failure shapes that recur in Spring Boot services under load, organised by layer. For each: what it looks like in the metrics from chapter 03, and the cheapest experiment that confirms or refutes it. Chapter 08 taught you to find the queue; this chapter is the field guide to what is usually standing in it.
One skill is needed throughout: seeing what PostgreSQL itself thinks happened. With pg_stat_statements enabled (chapter 05), ask the database:
SELECT calls, round(mean_exec_time::numeric, 2) AS mean_ms, round(total_exec_time::numeric, 1) AS total_ms, rows, left(query, 80) AS queryFROM pg_stat_statementsORDER BY total_exec_time DESCLIMIT 15;Two patterns in that output name most database problems before any other tool is opened: a query with enormous calls and small mean_ms (the N+1 shape — many cheap round trips) and a query with small calls and large mean_ms (an expensive plan — scan or sort). Both are in the catalog below.
And two actuator endpoints help enough to justify temporary exposure: threaddump for finding blocked threads and heapdump for leak hunts. Add them to the include list from chapter 03 for the duration of an investigation only — a heap dump contains live heap contents, and a thread dump is an attacker’s map of your internals. Neither belongs exposed in production.
Database bottlenecks
Missing index / full table scan. Metric signature: one endpoint’s server-side p99 rises first; pg_stat_statements shows its query with high mean_exec_time and rows scanned far above rows returned; DB-side CPU (not app CPU) climbs. Validate with EXPLAIN (ANALYZE, BUFFERS) on the exact query — Seq Scan on a million-row table is the answer. The trap: this never appears on small seed data (chapter 05), which is why the data volume matters.
N+1 queries. Metric signature: rate(spring_data_repository_invocations_seconds_count[1m]) runs at 10–50× the HTTP request rate; pg_stat_statements shows SELECT … FROM order_items WHERE order_id = ? with calls in the hundreds of thousands; HikariCP acquire time rises though each query is fast. Validate by counting queries per request — one page of 20 orders should issue 2 SQL statements (page + count), not 21. Chapter 10 breaks this on purpose so you can watch the signature appear.
Slow pagination. OFFSET n requires reading and discarding n rows; page 0 is cheap and page 5,000 is not. Signature: search p99 degrades with page number in the request mix, EXPLAIN shows Limit over a large Index Scan. Validate by paging to depth directly: curl '…/api/orders?page=5000'. The fix — keyset pagination — is chapter 10’s trade-off discussion, not the baseline.
Lock contention. Signature: pg_stat_activity shows sessions piled up in wait_event_type = 'Lock'; p99 jumps specifically on write endpoints while reads stay fast; throughput collapses though CPU is idle. Validate:
SELECT pid, state, wait_event_type, wait_event, left(query, 60) AS queryFROM pg_stat_activityWHERE wait_event_type = 'Lock';The usual causes: long transactions holding row locks, status updates hitting the same hot rows, or a migration lock. In the Order API, UPDATE on the same order_id from concurrent PATCHes serialises — realistic, and a reason write-heavy stress profiles exist.
Excessive transaction scope. A @Transactional wrapping external calls or long read-modify-write sequences holds its connection — and possibly locks — for the whole span. Signature: hikaricp_connections_usage_seconds p99 far exceeds the actual query time inside it. Validate by comparing connection-held time to query time; the fix is shrinking the transactional block, which chapter 10 demonstrates.
Application bottlenecks
Blocking I/O on the request thread. Any synchronous call — an HTTP call to another service, a file read, a Thread.sleep someone left — occupies a Tomcat thread for the full duration. Signature: tomcat_threads_busy approaches max while CPU stays low; thread dumps show many threads parked in the same socketRead. Validate with /actuator/threaddump and count threads in TIMED_WAITING/BLOCKED on one frame. (If you are on WebFlux rather than MVC, the equivalent sin is worse: a blocking call on an event-loop thread stalls every request the loop serves — never call blockable code there; offload to a bounded scheduler.)
Synchronised hotspots. synchronized collections or methods serialise threads. Signature: Tomcat busy climbs, CPU stays low, thread dumps show BLOCKED on one monitor. Validate the same way — the dump literally names the contended lock.
Expensive JSON serialisation. Deep or wide response objects spend CPU in Jackson and inflate payloads. Signature: process_cpu_usage scales with request rate while DB metrics stay flat; large Content-Length on responses. Validate by comparing request latency against orders_creation-style business timers — the gap is serialisation + mapping. The oversized-payload version of this (returning entities, megabyte responses) shows the same CPU shape plus client-side transfer time.
Allocation churn. Creating large intermediate objects per request — mapping through three DTO copies, String concatenation in loops — drives GC pressure rather than CPU-busy. Signature: rate(jvm_gc_memory_allocated_bytes_total) disproportionate to request rate, GC overhead climbing with load. Validate with the allocation-rate PromQL from chapter 08, then a profiler (async-profiler allocation mode) if confirmed.
Absent caching for hot reads. A read-heavy endpoint recomputing or re-querying identical data burns database CPU per request. Signature: pg_stat_statements shows the same query shape with calls proportional to requests and rows always small; latency is consistent but adds a floor. Validate by comparing request rate to query rate 1:1 — chapter 10 walks the caching change and, more importantly, its correctness cost.
Runtime bottlenecks
Undersized heap. Signature: GC events per second rising steeply with load (rate(jvm_gc_pause_seconds_count)), heap sawtooth compressed against max, overhead rate(jvm_gc_pause_seconds_sum[1m]) in double digits. Validate by checking whether the post-GC floor is near max — near-max floor plus frequent GC means the heap is genuinely too small, not just allocation-heavy. Raising -Xmx is legitimate here — it is the one case where “add resources” is the diagnosis, not a guess. But prove it first.
Allocation-rate-driven GC. Heap can be generously sized and GC still hot if the code allocates wastefully. Signature differs from the above: post-GC floor is low, but the sawtooth slope is steep — young generation fills constantly. Same metric family, different picture; the remedy is allocation reduction, not memory.
CPU throttling. On a host: uncommon. Under Kubernetes limits: the stealth bottleneck — the JVM sees 100% of its allowance while process_cpu_usage looks modest, and latency degrades in 100 ms CFS scheduling quanta. Signature: flat-ish process_cpu_usage + degraded latency + container_cpu_cfs_throttled_seconds_total rising. Chapter 12 covers this fully; it is listed here because misdiagnosing it as an application problem is classic.
Thread starvation. Not the Tomcat queue — the case where the pool itself is undersized for concurrent long-running work, or virtual threads would help (Java 21: spring.threads.virtual.enabled=true swaps Tomcat’s platform threads for virtual ones — needs validation for your workload; it changes the starvation signature rather than removing it).
Infrastructure bottlenecks
Container CPU/memory limits. Same throttling story; add OOM-kill restarts as the memory version — process_uptime_seconds resetting tells you the pod restarted, which is why uptime belongs on the dashboard.
Autoscaling lag. Signature: a spike test where errors cluster in the first N minutes, then service stabilises — HPA scaled, but the new pods’ JVMs warmed up during the burst. Chapter 12 separates “startup time” from “app slowness”; the metrics to compare are per-pod age versus per-pod latency.
Load balancer / ingress constraints. Connection limits, TLS termination CPU, keepalive settings between LB and pods. Signature: client-side latency ≫ server-side latency in the cluster (the chapter 08 gap test, run against per-pod metrics) while pod metrics stay flat.
Network saturation. Rare on localhost, real in cloud: bandwidth caps on instances, cross-AZ latency, conntrack exhaustion. Signature: client-server gap grows uniformly across endpoints with no corresponding saturation inside the app. Validate with the gap measurement itself — it is the only metric that isolates the network.
How to work the catalog
The catalog is a hypothesis generator, not a checklist to run top-to-bottom. The discipline stays chapter 07’s: one suspected layer, one experiment, one re-run. And the null result matters — if your signature does not match any row, that finding (the system fails in a way this catalog does not cover) is worth writing down on the run sheet, because it is exactly the kind of knowledge that does not survive undocumented.
Milestone check: for each layer — database, application, runtime, infrastructure — you can name one signature you could find in Grafana within sixty seconds and one SQL or actuator command that would confirm it. Chapter 10 takes the most common rows through their full fix cycle.