Gatling simulations
By the end of this chapter Gatling is wired into the Gradle build and you have four complete simulations covering the four workloads this series needs: read-heavy, mixed read/write, paginated search, and a stepped stress test that pushes past the SLOs from chapter 01. Along the way: feeders, session state, request correlation, the open-versus-closed model decision, and assertions that make a run fail loudly.
Add the Gatling plugin
Add one line to the plugins block of build.gradle.kts (chapter 03):
plugins { java id("org.springframework.boot") version "4.1.1" id("io.spring.dependency-management") version "1.1.7" id("io.gatling.gradle") version "3.15.1.3" // brings Gatling 3.15.1}The plugin adds a gatling source set — simulations live in src/gatling/java, test resources (feeder CSVs) in src/gatling/resources — and a gatlingRun task. With several simulation classes present, pick one by fully-qualified name (gatlingRun alone is interactive, and fails outright when CI=true):
./gradlew gatlingClasses # compile check only./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulation./gradlew gatlingRun --all # every simulation, sequentiallyGatling does not need the application as a dependency — it is a pure HTTP client. That separation is deliberate: the load generator must never share a JVM, a connection pool, or a class path with the system under test.
Open model, closed model — pick deliberately
Gatling’s injection API splits in two, and the choice determines what your numbers mean:
- Open model (
injectOpen): you control the arrival rate — new virtual users start at the configured rate regardless of whether earlier requests finished. If the server slows down, in-flight requests pile up, which is exactly what real traffic does. - Closed model (
injectClosed): you control concurrency — a fixed number of virtual users, each starting a new iteration only after finishing the last. If the server slows, the arrival rate drops automatically.
Principle: for a user-facing HTTP API, the open model is the honest one. Real visitors do not politely wait for earlier visitors to finish. The closed model hides saturation: as response times grow, the offered load silently shrinks to what the server can handle — this is coordinated omission, covered fully in chapter 07, and it is why a closed-model test can report “stable throughput” while every user is queued. Use closed models only when the real callers genuinely are bounded — a fixed fleet of batch workers, a thread pool of internal consumers — which is rare for REST APIs.
All four simulations below use the open model.
A note on feeders, headers, and auth
- Feeders inject per-request data into the virtual user’s session — the CSV exports from chapter 05.
circular()re-reads from the top when exhausted (right for read IDs);queue()consumes each row once and fails the run if exhausted (right forplaced_orders.csv— a PATCHed order is consumed, and reusing it would generate409s). - Session attributes are referenced in URLs and bodies with
#{name}— the CSV column name. - Headers belong on
HttpProtocolBuilderwhen they are universal (Accept,Content-Type), on the request otherwise. - Authentication: this lab API is unauthenticated. When yours is not, the pattern is a one-shot
execthat posts credentials and.check(jsonPath("$.token").saveAs("accessToken")), followed by.header("Authorization", "Bearer #{accessToken}")on each request — fetch tokens in abeforehook or once per user, never inside a measured request chain unless token issuance is itself in scope:
// Illustrative — not used by the lab API. Pattern for token-bearing APIs:// .exec(http("authenticate").post("/auth/token")// .body(StringBody("{\"client\":\"load-test\"}"))// .check(status().is(200), jsonPath("$.token").saveAs("accessToken")))// then on each request:// .header("Authorization", "Bearer #{accessToken}")Store real credentials in environment variables (System.getenv("LOADTEST_TOKEN")), never in the repository — a simulation with a committed token is a credential leak.
Simulation 1 — read-heavy
The dominant real-world shape: mostly fetches, some searches, no writes. Traffic distribution is 70% GET /api/orders/{id}, 30% search — deliberately unequal, because equal-weight endpoints are a lab fiction.
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;import io.gatling.javaapi.http.*;import java.time.Duration;
public class ReadHeavySimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http .baseUrl(System.getProperty("baseUrl", "http://localhost:8080")) .acceptHeader("application/json") .shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular(); FeederBuilder<String> customers = csv("data/customers.csv").circular();
ChainBuilder getOrder = feed(orders) .exec(http("GET /api/orders/{id}") .get("/api/orders/#{order_id}") .check(status().is(200)));
ChainBuilder searchOrders = feed(customers) .exec(http("GET /api/orders?customer") .get("/api/orders?customerId=#{customer_id}&size=20") .check(status().is(200)));
ScenarioBuilder readHeavy = scenario("read-heavy") .randomSwitch().on( percent(70.0).then(getOrder), percent(30.0).then(searchOrders));
{ setUp(readHeavy.injectOpen( // phase 1: warm-up — arrival rate ramps, results discarded rampUsersPerSec(1).to(30).during(Duration.ofMinutes(1)), // phase 2: steady state — the measurement window constantUsersPerSec(30).during(Duration.ofMinutes(3)), // phase 3: ramp-down — watches recovery, kept out of assertions rampUsersPerSec(30).to(0).during(Duration.ofSeconds(30)) ).protocols(httpProtocol)) .assertions( details("GET /api/orders/{id}").responseTime().percentile(99.0).lt(150), details("GET /api/orders?customer").responseTime().percentile(99.0).lt(300), global().failedRequests().percent().lt(0.1) ); }}shareConnections() makes virtual users share the underlying HTTP connection pool like real browsers and API clients behind keep-alive — without it, each virtual user gets a private connection and you benchmark TCP setup instead of your API. The assertions are the SLOs from chapter 01, expressed per request name: this run fails (non-zero exit, KO in the report) if p99 on fetches exceeds 150 ms.
Simulation 2 — mixed read/write
Adds create and status-update traffic in a realistic ratio. The chain that matters is createThenRead: the POST response’s id is saved into the session and immediately used in a follow-up GET — request correlation, the pattern for “create then fetch the thing you created”.
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;import io.gatling.javaapi.http.*;import java.time.Duration;
public class MixedWorkloadSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http .baseUrl(System.getProperty("baseUrl", "http://localhost:8080")) .acceptHeader("application/json") .contentTypeHeader("application/json") .shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular(); FeederBuilder<String> customers = csv("data/customers.csv").circular(); FeederBuilder<String> placed = csv("data/placed_orders.csv").queue();
ChainBuilder getOrder = feed(orders) .exec(http("GET /api/orders/{id}") .get("/api/orders/#{order_id}").check(status().is(200)));
ChainBuilder search = feed(customers) .exec(http("GET /api/orders?customer") .get("/api/orders?customerId=#{customer_id}&size=20") .check(status().is(200)));
ChainBuilder createThenRead = feed(customers) .exec(http("POST /api/orders") .post("/api/orders") .body(StringBody(""" {"customerId":#{customer_id}, "items":[{"sku":"SKU-LT-1","quantity":2,"unitPrice":19.99}, {"sku":"SKU-LT-7","quantity":1,"unitPrice":4.50}]} """)).asJson() .check(status().is(201), jsonPath("$.id").saveAs("createdOrderId"))) .pause(Duration.ofMillis(300)) .exec(http("GET created order") .get("/api/orders/#{createdOrderId}").check(status().is(200)));
ChainBuilder advanceStatus = feed(placed) .exec(http("PATCH /api/orders/{id}/status") .patch("/api/orders/#{order_id}/status") .body(StringBody("{\"status\":\"PAID\"}")).asJson() .check(status().is(200)));
ScenarioBuilder mixed = scenario("mixed").randomSwitch().on( percent(55.0).then(getOrder), percent(25.0).then(search), percent(12.0).then(createThenRead), percent(8.0).then(advanceStatus));
{ setUp(mixed.injectOpen( rampUsersPerSec(1).to(20).during(Duration.ofMinutes(1)), constantUsersPerSec(20).during(Duration.ofMinutes(3)) ).protocols(httpProtocol)) .assertions( global().responseTime().percentile(95.0).lt(250), global().failedRequests().percent().lt(0.5) ); }}Two details are load-bearing. First, placed uses queue() — each PLACED order can be transitioned to PAID exactly once, and if the feeder exhausts, the run fails rather than silently generating 409s. 12% of 20 users/s is 2.4 creations per second; the status-update pool (~120,000 PLACED orders) is far deeper than any single run consumes, which is why chapter 05’s template-reset strategy exists for repeated write runs. Second, every write is checked — a silent 500 on 8% of traffic would contaminate every latency percentile while looking like “the server held up”.
Why these weights? Example assumption: reads dominate e-commerce order traffic by roughly 5:1 in this fictional system. Your real ratio comes from production access logs — pull a day of traffic, count by route shape, and weight randomSwitch accordingly. The method is the deliverable, not the numbers.
Simulation 3 — paginated search
Search with pagination is where deep OFFSET and index quality show up. Each virtual user pages through three pages of results for one customer.
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;import io.gatling.javaapi.http.*;import java.time.Duration;
public class SearchSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http .baseUrl(System.getProperty("baseUrl", "http://localhost:8080")) .acceptHeader("application/json") .shareConnections();
FeederBuilder<String> customers = csv("data/customers.csv").circular();
ScenarioBuilder search = scenario("paginated search") .feed(customers) .exec(session -> session.set("page", 0)) .repeat(3).on( exec(http("GET /api/orders search page #{page}") .get(s -> "/api/orders?customerId=" + s.getString("customer_id") + "&page=" + s.getInt("page") + "&size=20") .check(status().is(200))) .pause(Duration.ofMillis(800)) .exec(session -> session.set("page", session.getInt("page") + 1)) );
{ setUp(search.injectOpen( rampUsersPerSec(1).to(15).during(Duration.ofMinutes(1)), constantUsersPerSec(15).during(Duration.ofMinutes(3)) ).protocols(httpProtocol)) .assertions( global().responseTime().percentile(99.0).lt(300), global().failedRequests().percent().lt(0.1) ); }}exec(session -> ...) is the escape hatch for session arithmetic — here, a page counter incremented per iteration. pause(800ms) models a human reading a page; in the open model it does not reduce the arrival rate, it only changes how long each virtual user stays active — which raises the concurrent-in-flight count by Little’s law. Both effects are intended.
Simulation 4 — stepped stress
This one’s job is to break the SLO on purpose, in controlled steps, so you can see where the knee is and what saturates first.
package in.o612.eng.orders.load;
import static io.gatling.javaapi.core.CoreDsl.*;import static io.gatling.javaapi.http.HttpDsl.*;
import io.gatling.javaapi.core.*;import io.gatling.javaapi.http.*;import java.time.Duration;
public class StressSimulation extends Simulation {
HttpProtocolBuilder httpProtocol = http .baseUrl(System.getProperty("baseUrl", "http://localhost:8080")) .acceptHeader("application/json") .contentTypeHeader("application/json") .shareConnections();
FeederBuilder<String> orders = csv("data/orders.csv").circular(); FeederBuilder<String> customers = csv("data/customers.csv").circular();
ScenarioBuilder probe = scenario("stepped stress").randomSwitch().on( percent(60.0).then(feed(orders).exec( http("GET /api/orders/{id}") .get("/api/orders/#{order_id}").check(status().is(200)))), percent(30.0).then(feed(customers).exec( http("GET /api/orders?customer") .get("/api/orders?customerId=#{customer_id}&size=20") .check(status().is(200)))), percent(10.0).then(feed(customers).exec( http("POST /api/orders") .post("/api/orders") .body(StringBody(""" {"customerId":#{customer_id}, "items":[{"sku":"SKU-ST","quantity":1,"unitPrice":9.99}]} """)).asJson() .check(status().is(201)))));
{ setUp(probe.injectOpen( // 25 → 150 req/s in six 2-minute steps, 30 s ramps between. // Widen the range until the SLO assertions fail; that level // IS the finding. incrementUsersPerSec(25.0) .times(6) .eachLevelLasting(Duration.ofMinutes(2)) .separatedByRampsLasting(Duration.ofSeconds(30)) .startingFrom(25.0) ).protocols(httpProtocol)) .assertions( global().responseTime().percentile(99.0).lt(300), global().failedRequests().percent().lt(0.5) ); }}The assertions here are not pass/fail gates — they are the tripwire that tells you which step crossed the line. Read the report’s “response time over time” graph against the injection profile: the step where p99 bends sharply upward is the knee, and the level just below it is the service’s honest capacity under this profile. Chapter 08 shows how to correlate that knee with the saturation metrics from chapter 03.
Injection profiles, briefly
| DSL | Shape | Use for |
|---|---|---|
rampUsersPerSec(a).to(b) | arrival rate linear a→b | warm-up, ramp-down |
constantUsersPerSec(r).during(t) | flat arrival rate | steady-state measurement |
incrementUsersPerSec(step).times(n).eachLevelLasting(t) | staircase | stress / capacity tests |
nothingFor(t) | silence | cool-down |
constantConcurrentUsers(n) / rampConcurrentUsers | fixed concurrency | closed model only — bounded caller pools |
atOnceUsers(n) | all at once | spike tests (chapter 11) |
Common failure: forgetting .protocols(httpProtocol) — the scenario runs against no HTTP config and fails cryptically. And naming requests ("GET /api/orders/{id}") is not cosmetic: the report and the details(...) assertions are keyed by those names, so use the route template, not a friendly sentence.
Milestone
./gradlew gatlingClasses # compiles./gradlew gatlingRun --simulation in.o612.eng.orders.load.ReadHeavySimulationA full run takes ~4.5 minutes including warm-up and ramp-down. The console ends with a path to an HTML report under build/reports/gatling/. Open it — chapter 07 is about making sure the numbers in it are trustworthy.