Production hardening and the final exercise: prove it, then ship it
Checkpoint tag: v1.0.0 — production-readiness checklist complete, known limitations documented, final scenario green end to end.
What will be built
The last chapter is a review, not a feature: the architecture re-inspected against the running system, the threat model updated from skeleton to as-built, four failure drills executed and recorded, secret rotation and data retention answered concretely, the upgrade strategy written down, and the final incident scenario — the one from Chapter 0 — run end to end as the acceptance test for the whole series.
Why it matters
Every prior chapter built a mechanism; this chapter answers “does the mechanism survive contact.” Readiness reviews are where the gap between implemented and operable shows up: a runbook nobody rehearsed, a dashboard nobody loads during the drill, a rotation procedure that requires downtime nobody noticed. You ship what you rehearsed.
The architecture review — as built
Walk the original boundaries and check they survived implementation:
| Boundary (Ch 0/1 promise) | As-built enforcement |
|---|---|
domain-contracts is framework-light | ArchUnit rule, still green; zero Spring AI imports |
| Model behind a port | ModelGateway; stub/Ollama/hosted interchangeable |
| Retrieval tenant-safe | WHERE tenant_id + supersede filter; eval case sec-001 guards it |
| Tool calls policy-gated | ToolPolicyRegistry in dispatch, not in prompts |
| Writes approval-gated | agent.approvals + arg-hash + single-use transition |
| Audit separated from logs | agent.audit_events append-only + AUDIT marker |
| Telemetry redacted | TelemetryRedactionTest + collector transform |
Any row you can’t point at in code is documentation debt — fix the code or fix the doc, never leave a phantom control.
Threat model — as-built update
Review the Chapter 0 table against shipped reality:
- Closed: cross-tenant retrieval/tool access (claim-vs-arg enforcement + eval cases), approval replay/TOCTOU (arg hash + conditional update), unbounded orchestration (three budget dimensions), telemetry leakage (double redaction).
- Residual, accepted: indirect injection can still produce a wrong answer — mitigated by citations, abstention, and the fact that answers can’t mutate anything; local-dev profiles relax controls (labeled); claims-based identity propagation trusts agent-api — token exchange remains an optional lab.
- Documented non-goals: multi-agent delegation, internet-facing deployment, real ITSM integration — the simulator stands in; the seams for the real thing are the contract types and
SimulatorClient.
Operational drills
Run each, record output in docs/runbooks/drill-results.md:
- Dependency death: kill Ollama →
MODEL_UNAVAILABLEfast-fails; trace shows the failure boundary; recovery on restart is clean (no stuck loops). - IdP outage: Keycloak down → 401s at both boundaries; deny metrics spike; recovery requires no code.
- Bad deploy rollback: deploy a deliberately broken
agent-apiimage tag → readiness fails → rollout halts (maxUnavailable: 0 did its job) →rollout undo. - Secret rotation: rotate
AGENT_MCP_CLIENT_SECRETin Keycloak, update Secret, rolling restart → zero failed auth on the old secret before expiry — because you rotate the Keycloak credential first and the Pod env second; the order is the lesson.
Retention, cost, and upgrade strategy
- Retention:
agent.messagesandaudit_eventsget stated policies (90d operational, per-org audit) with a purge job sketched; runbook chunks are self-cleaning via supersede. - Cost: token usage metrics from Chapter 6/12 roll up into a per-conversation estimate;
local-debugprompt capture stays file-only. - Upgrades: Spring AI and the MCP SDK move fast — the upgrade procedure is bump catalog →
./gradlew check→./gradlew :evaluation-suite:evaluate→ drill 1. The eval suite is the upgrade gate; that is why it exists.
The final exercise
One scenario, every mechanism:
priya(operator,acme) asks: “payment-gateway is returning 5xx since the last deploy — assess it, and open an incident.”
Expected sequence — each step verifiable in telemetry:
- Retrieval finds
rb-payment-gateway-degradedchunks (acme-scoped, current version). - The loop calls
get_service_status+list_recent_incidentsunderops:read—ToolProposedevents on SSE. IncidentAssessmentreturns schema-valid, evidence refs all real chunk IDs.create_incidentproposed → policy: HIGH risk,ops:incident:writepresent, approval required →ApprovalRequiredevent with args JSON and 5-min expiry.priyaapproves → arg-hash verified → idempotent execution → incidentinc-…created.- One Tempo trace contains the whole path;
agent.audit_eventsholdsAPPROVAL_REQUESTED,APPROVED,TOOL_EXECUTEDwith matching hashes. - The same question from
samdies at step 4 withINSUFFICIENT_SCOPE— and fromrinat retrieval, where globex can’t see acme’s runbook.
If all seven hold on a clean checkout, the series delivered what the blueprint promised.
Known limitations — v1.0.0
Single-node event bus and approval queue; claims-based identity propagation (not RFC 8693); simulator instead of real ITSM; eval judge coverage is thin by design; Compose is not HA; the K8s overlay assumes ingress+storage; model quality depends on the deployed model, which the eval suite measures but cannot guarantee.
Checkpoint verification checklist
- All four drills executed and recorded.
- Threat-model residual risks written down, not vibes.
- Final scenario green, including the two negative paths.
-
git tag v1.0.0on a tree where./gradlew clean check+evaluate+gatlingRunall pass.
What comes next
The optional labs — Java 27 structured concurrency, A2A, Kotlin ADK, reranking, GraalVM, local model benchmarks — each isolated, each safe to skip. What you have now is the thing most “AI tutorials” never produce: a system whose limits you can name because you built the boundaries yourself.
Project State Ledger — v1.0.0
- Complete: 9 modules, 5 MCP tools, 4 scopes, 2 tenants, full observability + eval + perf + deploy
- Verification suite:
check+evaluate+gatlingRun+ drill script - Series complete. Carry this ledger into any maintenance or extension session.