Series overview
Part 1 of 138% complete
2026-07-20•10 min read

A vocabulary for measuring

By the end of this chapter you can name the seven kinds of performance test, say which question each one answers, and write acceptance criteria for an API that a Gatling report can actually pass or fail. You also get the queueing model the rest of the series uses to explain why response times degrade non-linearly — and why a high requests-per-second number on its own proves nothing.

This series is for developers who already build and run Spring Boot services. You should be comfortable with Java, REST controllers, Spring Data, and basic unit or integration testing; none of those are explained here. What the series does explain is the measurement method: how to instrument a service, generate honest load, read the result, and change exactly one variable at a time. The running example is an Order Management API — create an order, fetch it, search orders by customer and status, update an order’s status — built on PostgreSQL and tested with Gatling.

Versions and labels used throughout the series

ComponentVersionCheck with
Java21java -version
Gradle9.x, via the wrapper./gradlew --version
Spring Boot4.1.1build.gradle.kts
PostgreSQL18SELECT version();
Gatling3.15.1 via io.gatling.gradle 3.15.1.3./gradlew dependencies
Prometheusv3.7.0docker inspect on the container
Grafana12.2.0docker inspect on the container
Docker Engine with Compose v2any current releasedocker compose version

Numbers in this series carry one of three labels, because a configuration that is right for one workload is wrong for another:

  • Principle: true of the JVM, Spring Boot, or queueing in general, with a link to the primary source where it matters.
  • Example assumption: a choice made for the Order Management API. Change it when your context differs.
  • Needs validation: a decision only your own measurements can settle. The series gives the method, not the answer.

Seven tests, seven questions

“Performance testing” is the umbrella term, not a test. Each variant answers a different question, generates a different traffic pattern, and fails in a different way. Conflating them is the most common reason a load-test report gets filed and ignored: the team ran a stress test, reported a throughput number, and nobody asked the question the number answers.

TestQuestion it answersTraffic patternExpected outcomeCommon failure signalExample use case
Performance testDoes the service meet its latency and throughput targets under expected load?Representative mix at target rate, held steadySLOs met for the full durationp95 or p99 above target while throughput looks fineRelease check: “p99 < 300 ms at 200 req/s”
Load testDoes the service hold up at the expected peak, for a sustained period?Peak business load, ramped then held for tens of minutesStable latency, no error growth, no resource exhaustionLatency creeping upward across the steady phasePre-launch validation of the Order API at projected Black Friday peak
Stress testWhere is the breaking point, and does the service recover?Load increased stepwise past capacity until SLOs failDegradation is gradual and recoverable, not a cliffErrors spike abruptly; recovery does not happen after load stopsFind the ceiling before marketing announces a campaign
Spike testWhat happens when traffic arrives far faster than the service or autoscaler can react?Sudden step change — e.g. 10x in under a minute — then back downQueueing and brief elevated latency, then return to SLOErrors during the spike; latency that never settles backFlash sale start, push-notification-driven burst
Soak (endurance) testDoes anything degrade over hours at moderate load?Moderate, steady load for 4–24+ hoursFlat memory, connection counts, and latencyHeap or connection pool climbing monotonically; GC pauses lengtheningCatching a connection leak that only matters over a weekend
Capacity testHow much load can one instance (or the cluster) serve within SLO?Incremental levels with a measurement at eachA defensible “max sustained req/s” numberThe knee: the level after which latency acceleratesSizing replicas; answering “how many pods for the launch?”
Scalability testDoes adding resources add proportional capacity?Same workload, repeated at 1, 2, 4 instancesNear-linear throughput gain until a shared bottleneck appearsThroughput flat or worse with more replicasProving the database, not the app, is the scaling limit

Two things in that table do more work than the rest. The traffic pattern column is the test design — a soak test with a spike profile is just a badly-run stress test. And the failure signal column is the exit criterion: you run a stress test specifically to observe the failure signal, so “the test produced errors” is the expected outcome, not a reason to stop early.

Example assumption: this series treats a “performance test” as the SLO-conformance check, “load test” as the sustained-peak variant, and keeps stress, spike, soak, capacity, and scalability as distinct profiles. Your organisation’s names may differ; the patterns and questions are what carry over.

Why high RPS alone is not success

A Gatling report’s headline number is throughput: requests per second the system completed. It is the least informative number on the page, for three reasons.

  1. Throughput is decided by the client, not the server. In an open-model test — the default mental model, and the right one for user-facing traffic — the arrival rate is what you configured. The server does not “achieve” 500 req/s; 500 req/s arrived, and the question is what happened to them.
  2. Completed requests can all be errors. A service that returns HTTP 500 in 2 ms will show spectacular throughput and latency. The assertions that matter are on response time percentiles and failure ratio, never on the count.
  3. RPS says nothing about tail latency. At high utilisation, queueing delay concentrates in the tail. Two runs can show identical mean latency and RPS while one has a p99 ten times worse — a difference that decides whether users in that tail abandon the checkout.

The success criterion for every test in this series is therefore an assertion bundle, not a throughput figure: latency percentiles per endpoint, a failed-request ratio, and — when relevant — a throughput floor. Chapter 06 shows these as Gatling assertions(...), so a run fails loudly instead of producing a report nobody reads.

The measurement vocabulary: SLI, SLO, SLA

Before you can say whether a test passed, you need the three terms that define “passed”:

  • SLI (service level indicator) — the measured quantity. For this API: request latency by endpoint, successful-response ratio, and throughput. An SLI must be measurable from instrumentation, not vibes.
  • SLO (service level objective) — the internal target on an SLI, over a window. “p99 latency on GET /api/orders/{id} stays below 300 ms over any 5-minute window” is an SLO. SLOs are what your load tests assert.
  • SLA (service level agreement) — the contractual promise to a customer, typically looser than the internal SLO and attached to penalties. The gap between SLA and SLO is your error budget — room to be imperfect without breaching the contract.

Two more terms do quiet but essential work:

  • Saturation — how full a resource is: CPU utilisation, HikariCP connections in use, Tomcat worker threads busy. Saturation is the cause you correlate latency against; chapter 08 is built around this.
  • Availability — the fraction of time (or requests) the service is usable, e.g. “99.9% of requests get a non-5xx response”. In load-test terms it maps to the failed-request ratio.

Percentiles, and why the average lies

Latency distributions under load are right-skewed: most requests are fast, a few are slow, and the slow ones are very slow. Reporting the mean of a skewed distribution describes a request nobody experienced.

  • p50 (median) — half of requests were faster. Use it for “typical” experience.
  • p95 — the request one user in twenty experiences. This is usually the first place overload shows.
  • p99 — one in a hundred. On an endpoint hit 500 times a second, p99 is exceeded five times every second — hundreds of real users per minute.
  • max — dominated by outliers: a single GC pause, a TCP retransmit, a container throttle slice. Useful as a tripwire (“max must stay under 2 s”), useless as a target.

The numerical example that makes this concrete: a service where p50 is 20 ms and p99 is 1.4 s has a mean somewhere around 40 ms. “Average response time 40 ms” reads as healthy and describes nothing real. Assert on p95 and p99, keep max as a guardrail, and treat mean as a smoke signal only.

Arrival rate, concurrency, and the queue

Three quantities control every load test, and they are related by a constraint you cannot configure away.

Arrival rate is how fast new requests arrive: “200 orders per second”. Concurrency is how many requests are in flight at once. Response time is how long each takes. Little’s law — a principle, not an approximation — ties them together:

concurrency = arrival rate × response time

At 200 req/s with a 100 ms response time, roughly 20 requests are in flight at any instant. Double the response time under load and concurrency doubles to hold the same arrival rate. This is why the relationship between load and latency is a knee, not a line: a system has finite servers (Tomcat worker threads, HikariCP connections, CPU cores). While arrival rate stays below service capacity, concurrency stays low and response time is flat. Once arrivals exceed what the servers can drain, requests queue, and every additional arrival adds queueing delay to everyone behind it.

Order API

Load generator

HTTP request

Gatling

accept queue

Tomcat worker threads

HikariCP connection pool

PostgreSQL

Order API

Load generator

HTTP request

Gatling

accept queue

Tomcat worker threads

HikariCP connection pool

PostgreSQL

A request’s total latency is the sum of time spent waiting at each stage plus the service time: accept queue → worker thread → connection pool → database → serialisation. Server-side latency is the portion inside the application and its dependencies; client-side latency — what Gatling measures — adds network transit and its own queueing. Chapters 07 and 08 keep these two strictly separate, because confusing them is how teams conclude “the API is slow” when the load generator’s own thread pool was saturated.

Principle: above the knee, latency is dominated by queueing, not by the work itself. That is why “add resources” is sometimes the fix and sometimes useless — if the queue is behind the database connection pool, doubling CPU does nothing. Diagnosing which queue filled is what chapter 09 is for.

Acceptance criteria for the Order Management API

Example assumption: the following targets are invented for the running example — plausible for a mid-traffic internal API, not derived from a real business requirement. Replace every number before reusing them.

Endpoint / propertySLO
GET /api/orders/{id}p99 < 150 ms, p50 < 25 ms
GET /api/orders?customerId=&status=p99 < 300 ms over any 5-minute window
POST /api/ordersp99 < 400 ms, p50 < 60 ms
PATCH /api/orders/{id}/statusp99 < 250 ms
Error rate (non-2xx, excluding 404s on bad IDs)< 0.1% of requests
Sustained targetthe above hold at 150 req/s mixed traffic for 30 minutes
Soakno upward trend in heap, GC pause time, or pool utilisation over 4 hours at 60% of target

Three properties make these testable rather than aspirational: every criterion names a metric, a threshold, and a window; the traffic target is attached to the latency targets, not separate; and the soak criterion is about trends, not values.

What chapter 02 builds

With the vocabulary settled, the next chapter creates the lab: a repository layout, a Docker Compose file for PostgreSQL, Prometheus, and Grafana, and — most important — the list of variables that must be controlled before any two test runs can be compared.

Spring BootJavaPerformanceTesting

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind