Series overview
Part 15 of 1788% complete
2026-09-07•4 min read

Test performance and resilience: measuring the framework, not the model

Checkpoint tag: chapter-14-performance — performance-tests runs reproducible Gatling scenarios, every reported number comes from a documented local run, and initial SLOs are drafted from measurements.

What will be built

A performance-tests module with Gatling scenarios for the REST message path, SSE concurrency, and read-tool throughput; a stub-model latency mode so application overhead is measurable without a real model; fault runs (latency, flaky, outage) quantifying how degradation propagates; a JVM-level look at virtual threads, pool sizing, and JFR; and docs/perf/results-template.md so every number in the repo is reproducible.

Why it matters

The single most common lie in AI-benchmark content is reporting “agent latency” that is 95% hosted-model time measured over hotel Wi-Fi. This chapter’s core discipline: measure the application’s overhead against a deterministic stub; measure model latency separately, labeled, and never conflate them. The stub isn’t a toy here — it’s the control that makes the framework’s cost visible.

Concepts explained

Gatling over k6 (ADR-009): JVM-native DSL, same toolchain, feeds results straight into the same observability stack. The scenarios are small on purpose — this is a smoke envelope, not a capacity plan.

Virtual threads under blocking MCP calls. The MCP client is synchronous by design; on virtual threads, blocking is cheap. The risk isn’t blocking — it’s pinning (synchronized blocks holding a carrier thread) and unbounded downstream concurrency. JFR’s jdk.VirtualThreadPinned events are how you check the first; the Chapter 11 bulkhead is the second.

Pool arithmetic. Tomcat threads, the HTTP client’s connection pool to MCP, the MCP server’s pool to the simulator, the JDBC pool — the slowest link’s concurrency is the system’s real limit. The scenario matrix exists to find that link rather than assume it.

Scenario matrix

ScenarioTargetAssertion
messageStubLoad50 rps for 60s, MODEL_PROVIDER=stubp95 app overhead within your measured budget; errors < 1%
sseFanout100 concurrent SSE subscribersno emitter leaks (memory flat after close)
readToolThroughputget_service_status through the loopper-tool p95 recorded; MCP→simulator hop cost isolated
degradedToolsame + simulator latency 2000tool timeout fires, loop completes or degrades, error shape correct
bulkheadSaturation20 concurrent vs. permit=8fast SATURATED, no queue blowup
realModelSmoke (opt-in)Ollama, 5 rps, 60smodel latency reported separately

The Gatling simulation (Java DSL) for the headline scenario:

performance-tests/src/gatling/java/in/o612/eng/opsagent/perf/MessageStubSimulation.java
package in.o612.eng.opsagent.perf;
import io.gatling.javaapi.core.*;
import io.gatling.javaapi.http.*;
import java.time.Duration;
import static io.gatling.javaapi.core.CoreDsl.*;
import static io.gatling.javaapi.http.HttpDsl.*;
public class MessageStubSimulation extends Simulation {
HttpProtocolBuilder http = http.baseUrl(System.getenv().getOrDefault(
"AGENT_URL", "http://localhost:8080"))
.header("Authorization", "Bearer " + System.getenv("PERF_TOKEN"));
ScenarioBuilder scn = scenario("message-stub")
.exec(http("create").post("/api/v1/conversations")
.check(jsonPath("$.conversationId").saveAs("convId")))
.exec(http("message").post("/api/v1/conversations/#{convId}/messages")
.body(StringBody("{\"message\":\"status check\"}"))
.check(status().is(200)));
{
setUp(scn.injectOpen(constantUsersPerSec(50).during(Duration.ofSeconds(60))))
.protocols(http)
.assertions(global().responseTime().percentile(95.0).lt(2000),
global().failedRequests().percent().lt(1.0));
}
}

Running and reporting

terminal
MODEL_PROVIDER=stub ./gradlew :agent-api:bootRun &
./gradlew :performance-tests:gatlingRun
# report -> performance-tests/build/reports/gatling/*/index.html

Every table in docs/perf/ records: machine (CPU/RAM), JVM flags, provider mode, dataset size (chunks in document_chunks — vector search latency means nothing without a corpus size), and the raw Gatling output path. If a number can’t name its run, it doesn’t ship.

JFR walkthrough: -XX:StartFlightRecording on agent-api during readToolThroughput; inspect jdk.VirtualThreadPinned (expect none — the path is lock-free by construction; find some and you found a bug worth a chapter footnote), jdk.SocketRead for MCP wait time, and allocation rate under the bulkhead.

SLO seeds from measurements

Drafted after the runs, not before: e.g., “p95 framework overhead per message < X ms at 50 rps on this hardware; tool-call hop adds < Y ms” — X and Y filled in by your run, kept honest by the template.

Failure-injection lab

degradedTool is the lab: compare readToolThroughput baseline vs. latency 2000 vs. flaky 0.5. The interesting observation is retry amplification — flaky at the simulator doubles call volume in the client (2 attempts); at 50 rps input that’s 100 rps downstream. Do the math before enabling retries anywhere bigger.

Checkpoint verification checklist

  • gatlingRun green on the stub profile with assertions enabled.
  • Real-model numbers labeled separately — no conflation in the report.
  • Bulkhead saturation tested, not assumed.
  • results-template.md filled for at least one run.

Commit message and Git tag

test(perf): Gatling scenarios, stub-model baselines, fault-mode runs, results template

git tag chapter-14-performance

What comes next

Chapter 15 packages everything — container images, the full Compose stack, and Kubernetes manifests that reflect what the previous fourteen chapters actually built.

Project State Ledger — chapter-14-performance

  • Module: performance-tests (Gatling Java DSL); scenarios: stub load, SSE fanout, tool throughput, degraded, saturation, opt-in real-model
  • Rules: stub isolates framework cost; model latency reported separately; every number cites its run
  • Diagnostics: JFR for VT pinning + socket waits; pool arithmetic documented
  • Next: chapter-15-deployment
PerformanceTestingJavaJVM

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind