Series overview
Part 12 of 1771% complete
2026-09-02•4 min read

Add guardrails and bounded orchestration: budgets, timeouts, and degraded modes

Checkpoint tag: chapter-11-guardrails-resilience — runaway loops, slow models, injected instructions, and retry storms all terminate safely and observably.

What will be built

The reliability layer: an explicit timeout hierarchy (HTTP → orchestration → per-tool → downstream), retry classification (retryable in the policy registry now means something), a concurrency bulkhead around model calls, prompt/context size enforcement end to end, safe handling of model-generated identifiers, and a DEGRADED response mode so a dead vector store or MCP server produces a bounded answer instead of a stack trace.

Why it matters

Every unbounded thing in an agent system is a billing vector or an outage vector. The model is slow; tools are slower; a retry storm multiplies both. The disciplines here are the same ones you’d apply to any fan-out service — budgets, deadlines, bulkheads — but the nondeterministic component makes them mandatory rather than nice: a model that loops is not a bug you fix, it’s a behavior you bound.

Concepts explained

Retry classification, not retry enthusiasm. The registry’s retryable flag exists because “503 from a status lookup” (retry, once, with jitter) and “403 tenant mismatch” (never retry) are different animals. Rules: retry only on transport errors and 5xx/429; never on 4xx, timeouts past deadline, or schema-invalid tool output; at most 2 attempts; jittered backoff (100–400 ms) so parallel callers don’t synchronize into a thundering herd.

The timeout hierarchy must compose inward. A 45s loop deadline containing 12s tool calls containing a 10s downstream read timeout: each layer’s budget must fit inside its parent’s, or outer timeouts silently truncate inner retries. The invariant is checked by a unit test, not by faith.

Bulkheads on the model. Semaphore around ModelGateway calls caps concurrent model invocations (default 8). Excess requests get Failed(code=SATURATED) immediately rather than queueing into memory. The model is the most expensive dependency in the system — protect its callers from each other first.

Degraded mode. RetrievalResult failure → answer without RAG (labeled “no runbook context available”); MCP unavailable → read tools return UPSTREAM_* results the model can summarize honestly; model itself down → Failed(code=MODEL_UNAVAILABLE) + 503. Each degradation is a product decision written in code, not an accident.

Files added or changed

agent-api/…/resilience/{RetryExecutor, DeadlineContext, ModelBulkhead, DegradedModeAdvisor}.java
agent-api/…/tools/McpToolExecutor.java (retry classification)
agent-api/…/config/ResilienceProperties.java
agent-api/src/test/... (resilience labs)

Complete code (the load-bearing parts)

agent-api/src/main/java/in/o612/eng/opsagent/agent/resilience/RetryExecutor.java
package in.o612.eng.opsagent.agent.resilience;
import java.util.concurrent.ThreadLocalRandom;
import java.util.function.Supplier;
public class RetryExecutor {
public static <T> T execute(Supplier<T> call, boolean retryable) {
int attempts = retryable ? 2 : 1;
RuntimeException last = null;
for (int i = 0; i < attempts; i++) {
try {
return call.get();
} catch (RuntimeException e) {
if (!isRetryable(e) || i == attempts - 1) throw e;
last = e;
sleep(100 + ThreadLocalRandom.current().nextInt(300)); // jitter
}
}
throw last;
}
static boolean isRetryable(RuntimeException e) {
// transport failures and 5xx/429 only; 4xx, auth, schema errors -> never
return switch (e) {
case org.springframework.web.client.ResourceAccessException rae -> true;
case org.springframework.web.client.HttpServerErrorException hse -> true;
default -> false;
};
}
private static void sleep(long ms) {
try { Thread.sleep(ms); } catch (InterruptedException ie) { Thread.currentThread().interrupt(); }
}
}
agent-api/src/main/java/in/o612/eng/opsagent/agent/resilience/ModelBulkhead.java
package in.o612.eng.opsagent.agent.resilience;
import org.springframework.stereotype.Component;
import java.util.concurrent.Semaphore;
import java.util.function.Supplier;
@Component
public class ModelBulkhead {
private final Semaphore permits = new Semaphore(8);
public <T> T guard(Supplier<T> call) {
if (!permits.tryAcquire()) {
throw new SaturatedException("model concurrency limit reached");
}
try { return call.get(); } finally { permits.release(); }
}
public static class SaturatedException extends RuntimeException {
public SaturatedException(String m) { super(m); }
}
}

The orchestrator wraps each chatClient…call() in bulkhead.guard(...); AgentOrchestrator.run passes its remaining deadline into executor.execute so a late-stage tool call gets its remaining budget, not a fresh 12 seconds — deadline propagation, not parallel clocks. McpToolExecutor wraps each callTool in RetryExecutor.execute(…, policy.retryable()).

Framework 7’s @Retryable annotation was the alternative; we chose the explicit executor because the retry decision needs the policy’s flag and our error classification, and Chapter 14 measures retry overhead you can’t see inside an annotation. Resilience4j stays on the bench — nothing here exceeds what two JDK classes express.

Failure-injection lab

The whole lab, mapped to observable outcomes:

InjectionExpected observable
model stub: infinite tool proposalsTOOL_BUDGET_EXCEEDED at cap; agent.loop.steps capped
simulator latency 13000TOOL_TIMEOUT at 12s; loop continues or degrades
simulator flaky 1.02 attempts, jittered, then UPSTREAM_* — count calls in the simulator log: exactly 2, not 10
20 concurrent requests with stub latency9th+ request gets SATURATED 503; no queue growth
50 KB promptPROMPT_TOO_LARGE pre-model; token spend = 0
Ollama downMODEL_UNAVAILABLE; retrieval still worked (degraded, not dead)

Security considerations

Bulkheads double as DoS containment; prompt caps double as cost containment. Tool results now pass through a size cap too (10 KB) — a compromised downstream can’t fill the context window with garbage. Model-proposed identifiers were already typed/bound; nothing in this chapter relaxes Chapter 9–10 gates — degraded mode narrows capability, never widens it.

Observability checks

agent.resilience.retries (tool × outcome), agent.resilience.saturated, agent.model.bulkhead.wait — plus the existing loop metrics. A retry-rate alert (rate(agent.resilience.retries[5m]) > threshold) is the early-warning for downstream sickness.

Checkpoint verification checklist

  • Timeout hierarchy composes (inner < outer) and is unit-tested as an invariant.
  • Retries: transport/5xx only, ≤2, jittered; 4xx never retried.
  • Saturation returns 503 fast, not slow.
  • Each dependency failure has a named degraded outcome.

Commit message and Git tag

feat(agent-api): bounded orchestration, retry classification, bulkheads, degraded modes

git tag chapter-11-guardrails-resilience

What comes next

Chapter 12 makes all of this visible: OpenTelemetry end to end, Prometheus metrics, Tempo traces, Loki logs, Grafana dashboards — with redaction, because the interesting payloads are exactly the ones you must not export.

Project State Ledger — chapter-11-guardrails-resilience

  • New: RetryExecutor (retryable={5xx,transport}, max 2, jitter), ModelBulkhead (8 permits, SATURATED), deadline propagation into per-tool calls, tool-result cap 10 KB
  • Degraded modes: RAG-down→labeled ungrounded answer; MCP-down→structured tool errors; model-down→503
  • Budgets: 8 model calls / 10 tool calls / 45s deadline (Ch 8), prompt ≤16k, context ≤12k chars
  • Decision: Framework-level retry annotation rejected — policy needs the flag + error classes; Resilience4j still unnecessary
  • Next: chapter-12-observability
JavaSpring BootAIPerformance

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind