Series overview
Part 3 of 1323% complete
2026-07-24•8 min read

Instrument before you measure

By the end of this chapter the Gradle project exists, the (still minimal) application starts, Prometheus scrapes it every five seconds, and you know which Micrometer metric answers each diagnosis question the rest of the series will ask. The rule being applied: instrument first. A load test against an unmeasured service produces a single number — client-side latency — and leaves every “why” unanswerable.

Create the Gradle project

Inside orders-perf-lab/, add settings.gradle.kts and build.gradle.kts, and generate the wrapper (gradle wrapper if you have Gradle installed, or copy the wrapper from any existing project):

settings.gradle.kts
rootProject.name = "orders-perf-lab"
build.gradle.kts
plugins {
java
id("org.springframework.boot") version "4.1.1"
id("io.spring.dependency-management") version "1.1.7"
}
group = "in.o612.eng"
version = "0.0.1-SNAPSHOT"
java {
toolchain {
languageVersion = JavaLanguageVersion.of(21)
}
}
repositories {
mavenCentral()
}
dependencies {
implementation("org.springframework.boot:spring-boot-starter-webmvc")
implementation("org.springframework.boot:spring-boot-starter-data-jpa")
implementation("org.springframework.boot:spring-boot-starter-validation")
implementation("org.springframework.boot:spring-boot-starter-actuator")
implementation("io.micrometer:micrometer-registry-prometheus")
runtimeOnly("org.postgresql:postgresql")
testImplementation("org.springframework.boot:spring-boot-starter-test")
testRuntimeOnly("org.junit.platform:junit-platform-launcher")
}
tasks.withType<Test> {
useJUnitPlatform()
}

Two notes on the choices:

  • spring-boot-starter-webmvc is the Spring Boot 4 name for the Spring MVC + embedded Tomcat starter (the old spring-boot-starter-web remains as a deprecated alias). Tomcat matters for this series: it is a thread-per-request server, so its worker-thread pool is one of the queues from chapter 01’s diagram.
  • micrometer-registry-prometheus is the registry that renders all Micrometer meters in the Prometheus text format at /actuator/prometheus. Nothing else is required — no annotations, no config class.

The application and its configuration

Add the entry point — deliberately minimal; the domain code arrives in chapter 04:

src/main/java/in/o612/eng/orders/OrderApiApplication.java
package in.o612.eng.orders;
import org.springframework.boot.SpringApplication;
import org.springframework.boot.autoconfigure.SpringBootApplication;
@SpringBootApplication
public class OrderApiApplication {
public static void main(String[] args) {
SpringApplication.run(OrderApiApplication.class, args);
}
}

Now application.yml, which does three jobs at once: connects to PostgreSQL, sets the JPA schema mode, and exposes the metrics endpoint:

src/main/resources/application.yml
spring:
application:
name: order-api
datasource:
url: jdbc:postgresql://localhost:5432/orders
username: orders
password: ${POSTGRES_PASSWORD:orders-local-pw}
hikari:
maximum-pool-size: 10
jpa:
hibernate:
ddl-auto: validate
open-in-view: false
server:
tomcat:
mbeanregistry:
enabled: true # required for tomcat_threads_* metrics (below)
threads:
max: 200 # the default, written down so it is a controlled variable
accept-count: 100
management:
endpoints:
web:
exposure:
include: health, info, prometheus, metrics
endpoint:
health:
probes:
enabled: true
prometheus:
metrics:
export:
enabled: true
metrics:
tags:
application: order-api
distribution:
percentiles-histogram:
http.server.requests: true
hikaricp.connections.acquire: true # bucketed wait-for-connection times
percentiles:
http.server.requests: 0.5, 0.95, 0.99
slo:
http.server.requests: 50ms, 100ms, 300ms

Read this file as a set of decisions, not boilerplate:

  • ddl-auto: validate — the schema is owned by db/init/01-schema.sql (chapter 04), and Hibernate merely verifies entities match it. update is convenient and ruins benchmarks silently: you stop knowing what your schema is.
  • open-in-view: false — Boot disables it by default; writing it down prevents a future “helpful” re-enable, which would hide lazy-loading queries inside the serialisation phase and make per-request query counts unmeasurable.
  • maximum-pool-size: 10 — HikariCP’s default, written down for the same reason. Needs validation: chapter 10 derives a size from measurements instead of accepting this.
  • percentiles-histogram: true — the important line. It makes http_server_requests publish Prometheus histogram buckets (_bucket series), which is what lets you compute p95/p99 in PromQL across any time window. Without it a Micrometer timer exports only _count/_sum/_max — the same flag on hikaricp.connections.acquire is what makes “how long did threads wait for a connection” answerable at p99 rather than at max. The percentiles list adds client-side-computed percentiles — useful for spot checks, but only histograms aggregate correctly across instances.
  • slo buckets — service-level-objective buckets aligned with chapter 01’s targets, so Grafana can show “fraction of requests inside SLO” directly.
  • mbeanregistry.enabled: true — registers Tomcat’s JMX MBeans, which is where tomcat_threads_* metrics come from. Without it, the Tomcat worker-thread section below reports nothing: Micrometer has no MBeans to read. This is a Spring Boot 4 behaviour worth knowing cold, because thread-pool saturation is invisible without it.
  • probes.enabled: true — publishes liveness/readiness; unused until Kubernetes in chapter 12, free to have now.

Security note: the exposure list is deliberately short — no heapdump, threaddump, or env, which leak memory contents and configuration. In anything beyond this lab, put the actuator endpoints on a separate management port or behind authentication; /actuator/prometheus reveals your internal metric names and label values to anyone who can reach it.

Milestone: metrics flowing

Terminal window
docker compose up -d # postgres must be healthy first
export POSTGRES_PASSWORD=orders-local-pw
./gradlew bootRun

Once Started OrderApiApplication appears:

Terminal window
curl -s http://localhost:8080/actuator/prometheus | head -20
curl -s http://localhost:8080/actuator/health

You should see # HELP/# TYPE lines for JVM metrics even though no HTTP traffic has happened yet — Micrometer starts recording at boot. Then open http://localhost:9090/targets: the order-api target should now be UP. If it is DOWN on Linux, the extra_hosts mapping from chapter 02 is missing; on Docker Desktop it works without it.

Until endpoints exist there is nothing under http_server_requests_seconds. The infrastructure is ready either way — that was the point of instrumenting first.

The metric catalog

Each group below names the Prometheus series Micrometer produces, the question it answers, and what a problem looks like. Chapter 08 turns these into a full query set; for now, learn the names.

HTTP request metrics

http_server_requests_seconds — a histogram (_count, _sum, _bucket) per (uri, method, status, outcome) combination. This is your server-side latency: measured inside the servlet filter chain, so it excludes network transit and any queueing that happened before the request reached the app. Gatling’s response times are client-side; the gap between the two is network plus client overhead, and chapter 07 uses exactly that gap to detect a saturated load generator.

Watch: p99 per uri, the ratio of status=~"5.." to total, and http_server_requests_seconds_count rate for achieved throughput as the server saw it.

JVM memory

jvm_memory_used_bytes{area="heap"} and {area="nonheap"}, plus jvm_memory_committed_bytes and jvm_memory_max_bytes. The heap series should show a sawtooth under load — rising to a peak, dropping at GC. A sawtooth whose floor keeps rising across a soak test is the classic leak signature (chapter 11). Non-heap growing without bound points at metaspace/direct-buffer issues, a different problem entirely.

Garbage collection

jvm_gc_pause_seconds (a summary: _count for GC event frequency, _sum for total pause time, tagged by action and cause). The derived quantity that matters is GC overhead: rate(jvm_gc_pause_seconds_sum[1m]) — fraction of wall-clock time spent stopped. Under a few percent, GC is healthy; sustained double digits means the JVM is collecting faster than the workload can afford. jvm_gc_memory_allocated_bytes_total gives the allocation rate driving it.

CPU and uptime

process_cpu_usage (the JVM’s share of CPU, where 1.0 means one fully-utilised core — a process using four cores reports 4.0), system_cpu_usage, and process_uptime_seconds. Correlate process CPU with latency: if p99 rises exactly when process_cpu_usage approaches the number of cores granted, you have a CPU-bound bottleneck. Under Kubernetes this metric reflects throttling — chapter 12 covers why that makes latency observations treacherous.

Threads and executors

jvm_threads_live_threads, jvm_threads_daemon_threads, and jvm_threads_states_threads{state="..."} — a request backlog shows up as growing runnable/blocked counts. Custom executor pools (a @Async executor, for instance) emit executor_* metrics when registered through Micrometer’s ExecutorServiceMetrics — which Spring Boot does automatically for auto-configured ThreadPoolTaskExecutor beans.

HikariCP

The connection-pool queue from chapter 01’s diagram, now with names:

  • hikaricp_connections_active / hikaricp_connections_idle / hikaricp_connections — pool state right now.
  • hikaricp_connections_pending — threads waiting for a connection. Any sustained non-zero value is a queue forming; this single series is the fastest confirmation of a pool-sized bottleneck.
  • hikaricp_connections_acquire_seconds — how long threads waited to get a connection.
  • hikaricp_connections_timeout_total — connections that timed out entirely; at the default 30 s timeout these are requests already lost.

Tomcat request threads

tomcat_threads_busy_threads and tomcat_threads_current_threads against server.tomcat.threads.max (200 here). When busy approaches max, new requests sit in the accept queue (server.tomcat.accept-count), and client-visible latency climbs while server-side metrics can still look calm — the requests have not reached the instrumentation yet. tomcat_connections_keepalive_* covers connection-level state.

Database-side visibility

Micrometer sees the pool, not the queries. For query-level evidence — slow statements, scans, lock waits — the series uses PostgreSQL’s own pg_stat_statements extension (enabled in chapter 05’s seed) and, where mentioned, EXPLAIN. Optional: the postgres_exporter container adds Postgres internals to Prometheus; the series does not require it, but in production you want it.

Custom business metrics

Framework metrics tell you that a request was slow; business timers tell you which part. Micrometer’s Timer wraps an operation; add this component now so chapter 04’s service can inject it:

src/main/java/in/o612/eng/orders/order/OrderMetrics.java
package in.o612.eng.orders.order;
import io.micrometer.core.instrument.Counter;
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Timer;
import org.springframework.stereotype.Component;
@Component
public class OrderMetrics {
private final Timer creationTimer;
private final Counter statusTransitions;
public OrderMetrics(MeterRegistry registry) {
this.creationTimer = Timer.builder("orders.creation")
.description("Time to persist a new order including its items")
.publishPercentileHistogram()
.register(registry);
this.statusTransitions = Counter.builder("orders.status.transitions")
.description("Successful order status changes")
.register(registry);
}
public Timer.Sample startCreationTimer() {
return Timer.start();
}
public void recordCreation(Timer.Sample sample) {
sample.stop(creationTimer);
}
public void countStatusTransition() {
statusTransitions.increment();
}
}

Two design points. First, meters are created once at bean construction — creating a Timer per request leaks meters and pollutes the registry. Second, publishPercentileHistogram() gives this timer the same bucket treatment as http_server_requests, so histogram_quantile works on it. Chapter 04 wires startCreationTimer()/recordCreation() around persistence, which lets you compare “time inside createOrder” against “total request time” — the difference being validation, mapping, and serialisation overhead.

Trade-off: every meter costs a little memory and CPU per recording. At hundreds of distinct label values (a timer tagged by customerId, say) cardinality explodes and Prometheus suffers. Keep labels to low-cardinality dimensions — endpoint, status, operation — never identifiers.

The full flow

Docker Compose

Order API - host

Load generator host

HTTP load

JDBC via HikariCP

scrape /actuator/prometheus every 5s

PromQL queries

HTML report

Gatling

Tomcat thread pool

OrderService

Micrometer registry

PostgreSQL

Prometheus

Grafana

results/

Docker Compose

Order API - host

Load generator host

HTTP load

JDBC via HikariCP

scrape /actuator/prometheus every 5s

PromQL queries

HTML report

Gatling

Tomcat thread pool

OrderService

Micrometer registry

PostgreSQL

Prometheus

Grafana

results/

The arrows are the causal chain for every diagnosis in this series: Gatling observes client-side latency; Micrometer exposes where server time went; Prometheus stores the history so a spike at 14:32 can be inspected at 15:00; Grafana renders it. Gatling’s HTML report and Grafana answer different questions — what happened versus why — and chapter 08 is built on reading them together.

Grafana dashboard checklist

Before the first test run, build (or plan) one dashboard covering exactly these panels. Resist dashboards with forty graphs — under load you have seconds to find the panel that matters:

  • Throughput: request rate per endpoint (rate(http_server_requests_seconds_count[1m]) by uri)
  • Latency: p50/p95/p99 per endpoint, from histogram buckets
  • Errors: 5xx and 4xx ratio, stacked by status
  • Saturation — CPU: process_cpu_usage vs cores available
  • Saturation — Tomcat: tomcat_threads_busy_threads vs max
  • Saturation — DB pool: active + pending HikariCP connections
  • JVM: heap used vs committed, GC pause rate, allocation rate
  • Postgres: connection count, and pg_stat_statements top queries if exported
  • Annotations: a marker at each test’s start/stop time — otherwise you cannot correlate the shape of a graph with the load profile that caused it

Milestone check: Prometheus shows order-api UP, curl returns metric lines, and you can name the series that would reveal a connection-pool queue (hikaricp_connections_pending) and a thread-pool queue (tomcat_threads_busy_threads at max). Chapter 04 builds the API these metrics will measure.

Spring BootJavaPerformanceTesting

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind