Performance tests in CI and Kubernetes
By the end of this chapter you have a CI job that runs a short performance regression on schedule, a policy for which test lives at which pipeline stage, and a Kubernetes manifest whose resource settings do not silently corrupt the measurements. The recurring theme: in CI and Kubernetes alike, the environment is part of the test — and an uncontrolled environment produces numbers that look like results.
A GitHub Actions regression job
A short, gated run — Gatling’s assertions are the pass/fail, so a broken SLO breaks the build. This runs a 5-minute profile, not a stress suite:
name: perf-regression
on: schedule: - cron: "0 3 * * *" # nightly — see the staging table below workflow_dispatch: # and on demand
jobs: perf: runs-on: ubuntu-latest timeout-minutes: 30
services: postgres: image: postgres:18-alpine env: POSTGRES_DB: orders POSTGRES_USER: orders POSTGRES_PASSWORD: ci-perf-pw ports: ["5432:5432"] options: >- --health-cmd "pg_isready -U orders -d orders" --health-interval 5s --health-retries 10 volumes: [] # see below: schema + seed applied by psql
steps: - uses: actions/checkout@v4
- uses: actions/setup-java@v4 with: distribution: temurin java-version: "21"
- uses: gradle/actions/setup-gradle@v4
- name: Initialise schema and seed data run: | export PGPASSWORD=ci-perf-pw psql -h localhost -U orders -d orders -f db/init/01-schema.sql psql -h localhost -U orders -d orders -f db/init/02-seed.sql
- name: Build and start the API run: | ./gradlew bootJar POSTGRES_PASSWORD=ci-perf-pw \ java -jar build/libs/orders-perf-lab-0.0.1-SNAPSHOT.jar & for i in $(seq 1 60); do curl -sf http://localhost:8080/actuator/health && break || sleep 2 done
- name: Run the regression simulation run: ./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation --no-daemon
- name: Archive the report if: always() uses: actions/upload-artifact@v4 with: name: gatling-report path: build/reports/gatling/ retention-days: 30Properties that matter: the job fails only via Gatling assertions (a failed SLO is a red build), the report uploads even on failure (if: always() — the report is the debugging material), and the whole thing is timed-boxed so a wedged run cannot hold a runner.
Two honest caveats. First, GitHub-hosted runners are shared VMs — treat CI numbers as a regression detector (did this commit move the needle?) and keep absolute SLO gates loose, because the machine underneath the test is not yours. Second, never point this at a shared environment by default; CI perf jobs that hit shared staging have cost teams real incidents.
Which test belongs at which stage
| Stage | Test | Duration | Gate |
|---|---|---|---|
| Pull request | None — or a 2–3 minute smoke profile at low rate, thresholds generous | ≤5 min | advisory only; PR runners are too noisy for strict p99 gates |
| Nightly build | Read-heavy + mixed at target rate; stepped stress weekly | 15–30 min | assertions fail the build; report archived |
| Staging | Full suite: load, stress, spike — against production-like data volume and limits | hours | release-blocking |
| Pre-release / pre-launch | Soak (4–24 h) + capacity test at expected peak | hours–day | sign-off required |
| Production | Never without authorization — and then only with rate limits and a rollback plan | — | — |
The pattern: the cheaper and noisier the environment, the earlier it runs and the looser its thresholds. A CI job’s job is catching “someone merged a query regression”; capacity truth comes from the staging run.
Kubernetes resources for a controlled benchmark
The Deployment below is sized for the lab’s target load — the numbers are an example assumption, and the point is the shape, not the values:
apiVersion: apps/v1kind: Deploymentmetadata: name: order-apispec: replicas: 2 selector: matchLabels: {app: order-api} template: metadata: labels: {app: order-api} spec: containers: - name: order-api image: registry.example.com/orders/order-api:1.0.0 ports: [{containerPort: 8080}] env: - name: POSTGRES_PASSWORD valueFrom: {secretKeyRef: {name: db-creds, key: password}} - name: SPRING_DATASOURCE_URL value: jdbc:postgresql://postgres:5432/orders resources: requests: cpu: "1" memory: 768Mi limits: memory: 768Mi # cpu limit deliberately omitted — see throttling section readinessProbe: httpGet: {path: /actuator/health/readiness, port: 8080} livenessProbe: httpGet: {path: /actuator/health/liveness, port: 8080} lifecycle: preStop: exec: {command: ["sh", "-c", "sleep 5"]} # let LB deregistration propagateapiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata: name: order-apispec: scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: order-api} minReplicas: 2 maxReplicas: 6 metrics: - type: Resource resource: name: cpu target: {type: Utilization, averageUtilization: 70}- Requests = scheduling guarantee.
requests.cpu: 1reserves a core’s worth of scheduling weight — during a benchmark, a pod that must share CPU time is a confound. Requests are the answer; limits are the problem (next section). - Memory limit equals request. Memory is incompressible — a pod over its limit is OOM-killed, so a limit matching request makes heap exhaustion a crash (visible, on the run sheet) rather than silent pressure.
- Probes wired to chapter 03’s
probes.enabled. Readiness gates traffic — essential for the HPA discussion below. preStopsleep. On scale-down or rollout, pod deletion is near-instant; a short grace lets endpoint removal propagate before the process exits, otherwise inflight requests die and look like application errors.
CPU throttling: the measurement hazard
This is the failure mode that makes Kubernetes benchmarks lie. A CPU limit is enforced as a CFS quota: with limits.cpu: 500m the container may use at most 50 ms of CPU per 100 ms period, on all cores combined. When a burst of requests needs more, the kernel freezes the container’s threads until the next period — in bursts of up to 100 ms at a time.
What that does to a latency measurement: p99 inflates by whole quanta while process_cpu_usage shows moderate utilisation — the JVM was prevented from being busy. The signature is container_cpu_cfs_throttled_seconds_total (and ..._throttled_periods_total) climbing during the run, on exactly the windows where p99 jumps.
rate(container_cpu_cfs_throttled_seconds_total{pod=~"order-api-.*"}[1m])Two consequences:
- For a controlled benchmark: omit CPU limits (requests still guarantee scheduling) or set them far above expected demand, and record the choice on the run sheet. Comparing runs under different limits is comparing different machines.
- For production realism: run with the production limits in place — the throttled behaviour is the production answer — but know that the number you report is “latency under these limits”, not “latency of this code”.
HPA, and the cold-pod problem
Horizontal scaling adds replicas, each of which boots a JVM: class-loading, JIT warm-up, HikariCP pool fill, shared_buffers population. For the first seconds-to-minutes, a new pod is measurably slower than a warm one — and the load balancer sends it traffic as soon as readiness passes.
So a spike test against an autoscaled deployment produces a characteristic picture: errors/latency during the scale-up lag, new pods slower than old pods for their first minutes, then convergence. Two disciplines keep the reading honest:
- Per-pod series. The
instancelabel from chapter 02’s Prometheus config separates pods; aggregate latency during a scale event blends cold pods with warm ones and hides the shape of the recovery. - “Time to healthy p99” is its own metric. From scale trigger to a pod matching the fleet’s p99 — that number, not the pod’s existence, is what autoscaling lag really costs. A short readiness ramp or a warm-up sidecar is a fix only if it shortens that interval measurably.
HPA also reacts on a metric loop (default ~15 s evaluation + scale-down stabilisation). For the spike profile in chapter 11, an autoscaler will usually be too slow to help inside the burst — which is itself a useful finding about your capacity margin.
Milestone check
You can now: run a gated regression in CI, place each test type at the right pipeline stage, deploy the API under Kubernetes with resource settings that don’t corrupt the measurement, and read an autoscale event without mistaking JVM warm-up for application slowness. Chapter 13 collects the whole method into the production checklist, the mistakes to avoid, and the capstone exercise.