Instrument before you measure
By the end of this chapter the Gradle project exists, the (still minimal) application starts, Prometheus scrapes it every five seconds, and you know which Micrometer metric answers each diagnosis question the rest of the series will ask. The rule being applied: instrument first. A load test against an unmeasured service produces a single number — client-side latency — and leaves every “why” unanswerable.
Create the Gradle project
Inside orders-perf-lab/, add settings.gradle.kts and build.gradle.kts, and generate the wrapper (gradle wrapper if you have Gradle installed, or copy the wrapper from any existing project):
rootProject.name = "orders-perf-lab"plugins { java id("org.springframework.boot") version "4.1.1" id("io.spring.dependency-management") version "1.1.7"}
group = "in.o612.eng"version = "0.0.1-SNAPSHOT"
java { toolchain { languageVersion = JavaLanguageVersion.of(21) }}
repositories { mavenCentral()}
dependencies { implementation("org.springframework.boot:spring-boot-starter-webmvc") implementation("org.springframework.boot:spring-boot-starter-data-jpa") implementation("org.springframework.boot:spring-boot-starter-validation") implementation("org.springframework.boot:spring-boot-starter-actuator") implementation("io.micrometer:micrometer-registry-prometheus") runtimeOnly("org.postgresql:postgresql")
testImplementation("org.springframework.boot:spring-boot-starter-test") testRuntimeOnly("org.junit.platform:junit-platform-launcher")}
tasks.withType<Test> { useJUnitPlatform()}Two notes on the choices:
spring-boot-starter-webmvcis the Spring Boot 4 name for the Spring MVC + embedded Tomcat starter (the oldspring-boot-starter-webremains as a deprecated alias). Tomcat matters for this series: it is a thread-per-request server, so its worker-thread pool is one of the queues from chapter 01’s diagram.micrometer-registry-prometheusis the registry that renders all Micrometer meters in the Prometheus text format at/actuator/prometheus. Nothing else is required — no annotations, no config class.
The application and its configuration
Add the entry point — deliberately minimal; the domain code arrives in chapter 04:
package in.o612.eng.orders;
import org.springframework.boot.SpringApplication;import org.springframework.boot.autoconfigure.SpringBootApplication;
@SpringBootApplicationpublic class OrderApiApplication {
public static void main(String[] args) { SpringApplication.run(OrderApiApplication.class, args); }}Now application.yml, which does three jobs at once: connects to PostgreSQL, sets the JPA schema mode, and exposes the metrics endpoint:
spring: application: name: order-api datasource: url: jdbc:postgresql://localhost:5432/orders username: orders password: ${POSTGRES_PASSWORD:orders-local-pw} hikari: maximum-pool-size: 10 jpa: hibernate: ddl-auto: validate open-in-view: false
server: tomcat: mbeanregistry: enabled: true # required for tomcat_threads_* metrics (below) threads: max: 200 # the default, written down so it is a controlled variable accept-count: 100
management: endpoints: web: exposure: include: health, info, prometheus, metrics endpoint: health: probes: enabled: true prometheus: metrics: export: enabled: true metrics: tags: application: order-api distribution: percentiles-histogram: http.server.requests: true hikaricp.connections.acquire: true # bucketed wait-for-connection times percentiles: http.server.requests: 0.5, 0.95, 0.99 slo: http.server.requests: 50ms, 100ms, 300msRead this file as a set of decisions, not boilerplate:
ddl-auto: validate— the schema is owned bydb/init/01-schema.sql(chapter 04), and Hibernate merely verifies entities match it.updateis convenient and ruins benchmarks silently: you stop knowing what your schema is.open-in-view: false— Boot disables it by default; writing it down prevents a future “helpful” re-enable, which would hide lazy-loading queries inside the serialisation phase and make per-request query counts unmeasurable.maximum-pool-size: 10— HikariCP’s default, written down for the same reason. Needs validation: chapter 10 derives a size from measurements instead of accepting this.percentiles-histogram: true— the important line. It makeshttp_server_requestspublish Prometheus histogram buckets (_bucketseries), which is what lets you compute p95/p99 in PromQL across any time window. Without it a Micrometer timer exports only_count/_sum/_max— the same flag onhikaricp.connections.acquireis what makes “how long did threads wait for a connection” answerable at p99 rather than at max. Thepercentileslist adds client-side-computed percentiles — useful for spot checks, but only histograms aggregate correctly across instances.slobuckets — service-level-objective buckets aligned with chapter 01’s targets, so Grafana can show “fraction of requests inside SLO” directly.mbeanregistry.enabled: true— registers Tomcat’s JMX MBeans, which is wheretomcat_threads_*metrics come from. Without it, the Tomcat worker-thread section below reports nothing: Micrometer has no MBeans to read. This is a Spring Boot 4 behaviour worth knowing cold, because thread-pool saturation is invisible without it.probes.enabled: true— publishes liveness/readiness; unused until Kubernetes in chapter 12, free to have now.
Security note: the exposure list is deliberately short — no heapdump, threaddump, or env, which leak memory contents and configuration. In anything beyond this lab, put the actuator endpoints on a separate management port or behind authentication; /actuator/prometheus reveals your internal metric names and label values to anyone who can reach it.
Milestone: metrics flowing
docker compose up -d # postgres must be healthy firstexport POSTGRES_PASSWORD=orders-local-pw./gradlew bootRunOnce Started OrderApiApplication appears:
curl -s http://localhost:8080/actuator/prometheus | head -20curl -s http://localhost:8080/actuator/healthYou should see # HELP/# TYPE lines for JVM metrics even though no HTTP traffic has happened yet — Micrometer starts recording at boot. Then open http://localhost:9090/targets: the order-api target should now be UP. If it is DOWN on Linux, the extra_hosts mapping from chapter 02 is missing; on Docker Desktop it works without it.
Until endpoints exist there is nothing under http_server_requests_seconds. The infrastructure is ready either way — that was the point of instrumenting first.
The metric catalog
Each group below names the Prometheus series Micrometer produces, the question it answers, and what a problem looks like. Chapter 08 turns these into a full query set; for now, learn the names.
HTTP request metrics
http_server_requests_seconds — a histogram (_count, _sum, _bucket) per (uri, method, status, outcome) combination. This is your server-side latency: measured inside the servlet filter chain, so it excludes network transit and any queueing that happened before the request reached the app. Gatling’s response times are client-side; the gap between the two is network plus client overhead, and chapter 07 uses exactly that gap to detect a saturated load generator.
Watch: p99 per uri, the ratio of status=~"5.." to total, and http_server_requests_seconds_count rate for achieved throughput as the server saw it.
JVM memory
jvm_memory_used_bytes{area="heap"} and {area="nonheap"}, plus jvm_memory_committed_bytes and jvm_memory_max_bytes. The heap series should show a sawtooth under load — rising to a peak, dropping at GC. A sawtooth whose floor keeps rising across a soak test is the classic leak signature (chapter 11). Non-heap growing without bound points at metaspace/direct-buffer issues, a different problem entirely.
Garbage collection
jvm_gc_pause_seconds (a summary: _count for GC event frequency, _sum for total pause time, tagged by action and cause). The derived quantity that matters is GC overhead: rate(jvm_gc_pause_seconds_sum[1m]) — fraction of wall-clock time spent stopped. Under a few percent, GC is healthy; sustained double digits means the JVM is collecting faster than the workload can afford. jvm_gc_memory_allocated_bytes_total gives the allocation rate driving it.
CPU and uptime
process_cpu_usage (the JVM’s share of CPU, where 1.0 means one fully-utilised core — a process using four cores reports 4.0), system_cpu_usage, and process_uptime_seconds. Correlate process CPU with latency: if p99 rises exactly when process_cpu_usage approaches the number of cores granted, you have a CPU-bound bottleneck. Under Kubernetes this metric reflects throttling — chapter 12 covers why that makes latency observations treacherous.
Threads and executors
jvm_threads_live_threads, jvm_threads_daemon_threads, and jvm_threads_states_threads{state="..."} — a request backlog shows up as growing runnable/blocked counts. Custom executor pools (a @Async executor, for instance) emit executor_* metrics when registered through Micrometer’s ExecutorServiceMetrics — which Spring Boot does automatically for auto-configured ThreadPoolTaskExecutor beans.
HikariCP
The connection-pool queue from chapter 01’s diagram, now with names:
hikaricp_connections_active/hikaricp_connections_idle/hikaricp_connections— pool state right now.hikaricp_connections_pending— threads waiting for a connection. Any sustained non-zero value is a queue forming; this single series is the fastest confirmation of a pool-sized bottleneck.hikaricp_connections_acquire_seconds— how long threads waited to get a connection.hikaricp_connections_timeout_total— connections that timed out entirely; at the default 30 s timeout these are requests already lost.
Tomcat request threads
tomcat_threads_busy_threads and tomcat_threads_current_threads against server.tomcat.threads.max (200 here). When busy approaches max, new requests sit in the accept queue (server.tomcat.accept-count), and client-visible latency climbs while server-side metrics can still look calm — the requests have not reached the instrumentation yet. tomcat_connections_keepalive_* covers connection-level state.
Database-side visibility
Micrometer sees the pool, not the queries. For query-level evidence — slow statements, scans, lock waits — the series uses PostgreSQL’s own pg_stat_statements extension (enabled in chapter 05’s seed) and, where mentioned, EXPLAIN. Optional: the postgres_exporter container adds Postgres internals to Prometheus; the series does not require it, but in production you want it.
Custom business metrics
Framework metrics tell you that a request was slow; business timers tell you which part. Micrometer’s Timer wraps an operation; add this component now so chapter 04’s service can inject it:
package in.o612.eng.orders.order;
import io.micrometer.core.instrument.Counter;import io.micrometer.core.instrument.MeterRegistry;import io.micrometer.core.instrument.Timer;import org.springframework.stereotype.Component;
@Componentpublic class OrderMetrics {
private final Timer creationTimer; private final Counter statusTransitions;
public OrderMetrics(MeterRegistry registry) { this.creationTimer = Timer.builder("orders.creation") .description("Time to persist a new order including its items") .publishPercentileHistogram() .register(registry); this.statusTransitions = Counter.builder("orders.status.transitions") .description("Successful order status changes") .register(registry); }
public Timer.Sample startCreationTimer() { return Timer.start(); }
public void recordCreation(Timer.Sample sample) { sample.stop(creationTimer); }
public void countStatusTransition() { statusTransitions.increment(); }}Two design points. First, meters are created once at bean construction — creating a Timer per request leaks meters and pollutes the registry. Second, publishPercentileHistogram() gives this timer the same bucket treatment as http_server_requests, so histogram_quantile works on it. Chapter 04 wires startCreationTimer()/recordCreation() around persistence, which lets you compare “time inside createOrder” against “total request time” — the difference being validation, mapping, and serialisation overhead.
Trade-off: every meter costs a little memory and CPU per recording. At hundreds of distinct label values (a timer tagged by customerId, say) cardinality explodes and Prometheus suffers. Keep labels to low-cardinality dimensions — endpoint, status, operation — never identifiers.
The full flow
The arrows are the causal chain for every diagnosis in this series: Gatling observes client-side latency; Micrometer exposes where server time went; Prometheus stores the history so a spike at 14:32 can be inspected at 15:00; Grafana renders it. Gatling’s HTML report and Grafana answer different questions — what happened versus why — and chapter 08 is built on reading them together.
Grafana dashboard checklist
Before the first test run, build (or plan) one dashboard covering exactly these panels. Resist dashboards with forty graphs — under load you have seconds to find the panel that matters:
- Throughput: request rate per endpoint (
rate(http_server_requests_seconds_count[1m])byuri) - Latency: p50/p95/p99 per endpoint, from histogram buckets
- Errors: 5xx and 4xx ratio, stacked by status
- Saturation — CPU:
process_cpu_usagevs cores available - Saturation — Tomcat:
tomcat_threads_busy_threadsvsmax - Saturation — DB pool: active + pending HikariCP connections
- JVM: heap used vs committed, GC pause rate, allocation rate
- Postgres: connection count, and
pg_stat_statementstop queries if exported - Annotations: a marker at each test’s start/stop time — otherwise you cannot correlate the shape of a graph with the load profile that caused it
Milestone check: Prometheus shows order-api UP, curl returns metric lines, and you can name the series that would reveal a connection-pool queue (hikaricp_connections_pending) and a thread-pool queue (tomcat_threads_busy_threads at max). Chapter 04 builds the API these metrics will measure.