Series overview
Part 8 of 1362% complete
2026-08-01•6 min read

Reading the results

By the end of this chapter you can open a Gatling report, decide in under a minute whether the run is worth analysing, and walk from a bent latency curve to a ranked list of suspects using Prometheus. The output artifact is a troubleshooting table and a decision tree you will reuse in chapters 09–11.

Anatomy of a Gatling report

Each run writes a self-contained HTML report under build/reports/gatling/<simulation>-<timestamp>/index.html. In reading order:

  1. Assertions banner — OK/KO against the assertions(...) you set. A KO is a finding, not a build artifact to shrug at: it names the exact request group that crossed its threshold.
  2. Global statistics table — per request name: count, min, max, mean, stdDev, percentiles, req/s, and error percentage. Read the error percentage first — a slow run is diagnosable, a failed run needs the failure understood before latency means anything.
  3. Response time percentiles over time — the single most useful chart. A flat line is health; a knee is capacity; a sawtooth correlated with GC is a JVM story; a staircase that tracks incrementUsersPerSec steps tells you exactly which load level broke the SLO.
  4. Response time distribution — the histogram shape. A tight low-latency mode plus a long tail is normal; a second hump at high latency means two populations of requests — usually “the ones that hit the cache/index” and “the ones that did not”.
  5. Active users along the run — in the open model this climbs during queueing (users arrive faster than they drain). A growing active-user curve is queued load rendered visible.
  6. Requests/responses per second — achieved throughput. Compare the offered rate (your injection profile) to the completed rate: divergence after some timestamp is the moment the service stopped keeping up.

Reading order that works: assertions → errors → percentile-over-time → the timestamp where it bent → that same timestamp in Grafana.

The PromQL query set

These are the queries behind the dashboard checklist from chapter 03. All assume the application="order-api" label from chapter 02’s scrape config.

Latency percentiles per endpoint (the server-side counterpart of Gatling’s client-side numbers):

histogram_quantile(0.99,
sum by (le, uri) (
rate(http_server_requests_seconds_bucket{application="order-api"}[1m])))

Error ratio (5xx share of all requests):

sum(rate(http_server_requests_seconds_count{application="order-api", status=~"5.."}[1m]))
/
sum(rate(http_server_requests_seconds_count{application="order-api"}[1m]))

Throughput as the server saw it:

sum by (uri) (rate(http_server_requests_seconds_count{application="order-api"}[1m]))

Heap utilisation and trend:

sum(jvm_memory_used_bytes{application="order-api", area="heap"})
/
sum(jvm_memory_max_bytes{application="order-api", area="heap"})

GC pause pressure — fraction of the last minute spent stopped, plus frequency:

sum(rate(jvm_gc_pause_seconds_sum{application="order-api"}[1m])) -- overhead ratio
sum(rate(jvm_gc_pause_seconds_count{application="order-api"}[1m])) -- pauses per second

Allocation rate (what drives GC):

rate(jvm_gc_memory_allocated_bytes_total{application="order-api"}[1m])

HikariCP pool utilisation and queueing:

hikaricp_connections_active{application="order-api"}
/ hikaricp_connections_max{application="order-api"} -- utilisation
hikaricp_connections_pending{application="order-api"} -- queued threads; should be 0
histogram_quantile(0.99,
sum by (le) (rate(hikaricp_connections_acquire_seconds_bucket{application="order-api"}[1m])))
-- wait time for a connection

Tomcat worker saturation:

tomcat_threads_busy_threads{application="order-api"}
/ tomcat_threads_config_max_threads{application="order-api"}

Process CPU:

process_cpu_usage{application="order-api"} -- 1.0 ≈ one core saturated

Two PromQL notes that bite in practice: always rate the _bucket/_count/_sum series, never the raw counters, and pick a window ([1m]) shorter than your load-test phases so the ramp-up does not smear into the steady state. If a query returns nothing, check up{job="order-api"} first — the target may be DOWN, and Grafana shows silence, not an error.

The correlation method

Numbers become a diagnosis when two views of the same second disagree or agree:

  1. In the Gatling report, find the timestamp where p99 first bends (or errors begin).
  2. In Grafana, pin the same timestamp. Which saturation metric moved first — CPU, Tomcat threads, HikariCP pending, GC?
  3. The metric that moved at or just before the bend is the primary suspect; metrics that moved after are symptoms of the pile-up behind it.
  4. Form one hypothesis (“connection pool exhausted at 40 req/s”), one experiment (chapter 09/10), one re-run.

The pairs that matter most:

Gatling showsGrafana showsReading
p99 climbshikaricp_connections_pending > 0 at the same timeDB-pool queue — requests wait for connections
p99 climbstomcat_threads_busy / max ≈ 1Tomcat queue — work is waiting for a worker thread
p99 climbsprocess_cpu_usage at the granted-core ceilingCPU-bound
p99 climbsGC overhead rising into double digitsGC-bound — usually allocation rate, not “bad GC”
p99 climbsall of the above flatLook downstream — the database, or the client itself
p99 flaterrors risingFailure path is fast — check status codes before latency
client p99 climbsserver p99 flatGenerator or network — the measurement, not the service

Troubleshooting table

The working artifact of this chapter. “First safe experiment” means the cheapest change that would disprove the hypothesis — falsifiable, reversible.

SymptomLikely bottleneckEvidence to collectFirst safe experimentCommon incorrect conclusion
p99 rises at a load level; process_cpu_usage at ceilingCPU saturationGC overhead, per-endpoint split, allocation rateHalve per-request work (smaller page or payload) and re-run — if the knee moves proportionally, CPU confirmed“Add more instances” — correct only if the bottleneck is per-instance CPU, wrong if the shared DB is already hot
p99 rises; hikaricp_connections_pending climbs from 0Connection pool exhaustionacquire_seconds p99, query count per request, pg_stat_statements top queriesFix the query count first (N+1 check, chapter 10); only then resize the pool“Raise maximum-pool-size” — moves the queue into PostgreSQL’s max_connections or worse, onto a CPU-bound DB
p99 rises; Tomcat busy threads pinned at max, CPU and DB calmWorker-thread starvation — something inside requests blocksThread dump or threaddump actuator; look for BLOCKED/TIMED_WAITING on one frameFind and bound the blocking call (timeout, async boundary); re-run“Raise server.tomcat.threads.max” — adds queue depth, not throughput; the blocker still caps you
Long tail of very slow requests while p50 is finePeriodic stall: GC pause, disk flush, lockjvm_gc_pause max vs p99 gap, pg_stat_activity waitsCorrelate the tail timestamps with GC events“Average latency is fine, ship it” — the tail is the user experience at this percentile
Latency and errors climb over minutes, never recover during runResource leak: pool, memory, file handlesHeap floor after GC, hikaricp_connections total, open socketsSoak test with heap dumps at intervals (chapter 11)“Raise the heap/limits” — postpones the failure off the end of the test, not out of production
Everything slow from the first request, all saturation flatCold caches / wrong environmentRun sheet variables, first-minute vs steady-state splitExtend warm-up; re-run“The code got slower” — you measured a cold JVM
Client p99 ≫ server p99, gap grows with rateLoad generator saturatedGenerator CPU, its own event-loop metricsRe-run from a bigger box or two generators“API regression” — the most expensive wrong conclusion available; it sends everyone hunting a defect that does not exist
Error rate climbs while latency stays lowFast failure path — timeouts, 429s, pool timeout_totalResponse code histogram, hikaricp_connections_timeout_totalRead the actual status codes; a 500 in 3 ms is a different problem than a 30 s timeout“Overloaded” — fast errors usually mean a limit was hit, not exceeded gradually

The decision tree

no: client–server gap grew

yes

yes

no

yes

no

yes

yes

no

no

p99 degraded in steady state

Server p99 also degraded?

Generator or network

check load-gen CPU

hikaricp_connections_pending > 0?

DB pool queue

→ chapter 10: query count / pool sizing

tomcat busy threads at max?

Worker starvation

thread dump → find the block

process_cpu at ceiling?

GC overhead high?

Allocation/GC bound

→ chapter 9: reduce allocation

CPU-bound work

profile the hot path

Downstream / DB server

pg_stat_statements, DB host metrics

no: client–server gap grew

yes

yes

no

yes

no

yes

yes

no

no

p99 degraded in steady state

Server p99 also degraded?

Generator or network

check load-gen CPU

hikaricp_connections_pending > 0?

DB pool queue

→ chapter 10: query count / pool sizing

tomcat busy threads at max?

Worker starvation

thread dump → find the block

process_cpu at ceiling?

GC overhead high?

Allocation/GC bound

→ chapter 9: reduce allocation

CPU-bound work

profile the hot path

Downstream / DB server

pg_stat_statements, DB host metrics

The tree encodes one rule: queues are found by looking at what is waiting, not at what is busy. A system can be 100% busy on CPU and healthy; pending > 0 on a connection pool is never healthy under target load.

Milestone check

You can now take any Gatling report, locate the moment the SLO broke, and produce — from Prometheus alone — a one-sentence hypothesis with named evidence. Chapter 09 is the catalog those hypotheses come from: the recurring bottleneck shapes, each with its metric signature and the experiment that proves it.

Spring BootJavaPerformanceTesting

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind