Series overview
Part 1 of 176% complete
2026-08-14•9 min read

What we are building: an enterprise operations agent

Checkpoint tag: chapter-00-architecture — this chapter adds documentation only; there is nothing to compile yet.

Assumptions for this chapter and the rest of the series: Java 25, Spring Boot 4.1.1, Spring AI 2.0.1, Kotlin 2.3.21, Gradle 9.7.1, PostgreSQL 17 with pgvector 0.8.6, Keycloak 26.7.x, and Docker Compose. Versions were verified against official release pages when this blueprint was written; each chapter re-states what it depends on. The fictional domain — an internal operations platform for two tenants, acme and globex — recurs across the series, so the boundaries drawn here are the seams the code cuts along later.

What will be built

By the last chapter you will have ops-agent-platform, a multi-module Gradle repository containing four deployable services and a verification suite. An ops engineer sends a question to agent-api; the agent retrieves relevant runbook sections with citations, optionally calls read-only MCP tools for live service status and incident data, proposes mutating tools only when policy allows, and executes writes only after a human approves a single-use token bound to the exact arguments. Every step lands in one distributed trace.

This chapter lays out the whole map first: what the system does, what it deliberately refuses to do, where the trust boundaries sit, and which versions and ADRs everything downstream depends on. Skipping this chapter to reach the code sooner is the usual reason agent projects end up as a chatbot with a YAML file glued on.

Why it matters

Most JVM-side AI material stops at “call the model, print the answer.” Production stops much later: after you have bounded the agent loop, decided who may invoke which tool, proven that a mutated argument cannot slip past an approval, and can show an auditor a trace for the incident ticket the agent filed at 03:00. The hard parts are the same engineering disciplines as any distributed system — isolation, authorization, idempotency, observability — applied to a component that is nondeterministic by design. This series treats the LLM as one unreliable dependency inside a normal Spring system, not as the system.

Prerequisites and starting point

There is no starting Git tag — this chapter creates the documentation skeleton. You need nothing installed. To follow the series as a whole you should be comfortable with Java, Spring Boot, REST, SQL, Docker, and Gradle; you do not need prior agent, MCP, or Kotlin experience.

Chatbot, RAG app, tool-using agent, MCP server, multi-agent system

These terms get used interchangeably in marketing copy; in an architecture document they are different systems with different failure modes.

  • Chatbot: a model plus a prompt template. Stateless-ish text in, text out. Failure mode: it says something wrong confidently.
  • RAG application: a chatbot plus retrieval. Answers are grounded in documents you control, which adds a new failure mode — retrieving the wrong document, or a document a tenant was never meant to see.
  • Tool-using agent: a model that can propose structured calls against real systems. New failure modes: wrong tool, wrong arguments, unbounded loops, and — the dangerous one — a correct-looking call nobody authorized.
  • MCP server: not an agent at all. It is a capability boundary — a service that exposes tools, resources, and prompts over a protocol so that any compliant client can discover and invoke them. Putting operational capability behind MCP instead of local @Tool beans is what gives us an independent authorization boundary.
  • Multi-agent system: agents delegating to agents. This series deliberately stops short of that (there is an optional A2A lab) because every additional delegating hop multiplies the audit and authorization surface.

ops-agent-platform is a tool-using agent talking to a remote MCP server, with RAG on one side and an approval gate on the other.

System context and architecture

Supporting services

ops-agent-platform

HTTPS + JWT

query + embeddings

chunks + embeddings

Streamable HTTP + JWT

REST

chat + embeddings

JWT validation

JWT validation

OTLP

OTLP

OTLP

OTLP

dashboards, audit events

Ops engineer

Auditor / on-call reviewer

agent-api

mcp-operations-server

operations-simulator

knowledge-ingestion

PostgreSQL + pgvector

Keycloak

Ollama or hosted model

OTel Collector + Prometheus + Tempo + Loki + Grafana

Supporting services

ops-agent-platform

HTTPS + JWT

query + embeddings

chunks + embeddings

Streamable HTTP + JWT

REST

chat + embeddings

JWT validation

JWT validation

OTLP

OTLP

OTLP

OTLP

dashboards, audit events

Ops engineer

Auditor / on-call reviewer

agent-api

mcp-operations-server

operations-simulator

knowledge-ingestion

PostgreSQL + pgvector

Keycloak

Ollama or hosted model

OTel Collector + Prometheus + Tempo + Loki + Grafana

The simulator stands in for the real systems of record — a service catalog, deployment registry, and incident tracker — so the series is self-contained and can inject failures on demand. In a real deployment, mcp-operations-server would wrap ServiceNow, PagerDuty, or your internal equivalents; the agent code would not change.

Module boundaries

Verification

Libraries

Applications

MCP Streamable HTTP

REST

testOnly

testOnly

testOnly

agent-api

knowledge-ingestion

mcp-operations-server

operations-simulator

domain-contracts

test-support

architecture-tests

evaluation-suite

performance-tests

Verification

Libraries

Applications

MCP Streamable HTTP

REST

testOnly

testOnly

testOnly

agent-api

knowledge-ingestion

mcp-operations-server

operations-simulator

domain-contracts

test-support

architecture-tests

evaluation-suite

performance-tests

The rule that matters most: domain-contracts — identifiers, commands, results, the IncidentAssessment type, error envelopes — depends on nothing from Spring AI, JPA, HTTP, or a model provider. AI types may flow inward through ports; they may not leak into contracts. ArchUnit enforces this from Chapter 1.

The two request shapes

A read-only question (RAG) never touches a tool:

Chat modelpgvectoragent-apiOps engineerChat modelpgvectoragent-apiOps engineerPOST /conversations/{id}/messagesauthenticate + authorize + policysimilarity search, tenant filter, thresholdchunks + metadatasystem prompt + delimited context + questiongrounded answeranswer + citations (or abstention)
Chat modelpgvectoragent-apiOps engineerChat modelpgvectoragent-apiOps engineerPOST /conversations/{id}/messagesauthenticate + authorize + policysimilarity search, tenant filter, thresholdchunks + metadatasystem prompt + delimited context + questiongrounded answeranswer + citations (or abstention)

A read-only tool call adds the MCP hop:

Loperations-simulatormcp-operations-serverTool policyagent-apiOps engineerLoperations-simulatormcp-operations-serverTool policyagent-apiOps engineerwhat is the status of payment-gatewaymodel proposes get_service_statustool call proposalis this tool allowed for this callerallowed, low risk, no approval neededMCP call + JWT + tenant contextGET /sim/v1/services/{id}/healthhealth payloadtyped tool resulttool result appended to contextfinal answer citing the toolanswer
Loperations-simulatormcp-operations-serverTool policyagent-apiOps engineerLoperations-simulatormcp-operations-serverTool policyagent-apiOps engineerwhat is the status of payment-gatewaymodel proposes get_service_statustool call proposalis this tool allowed for this callerallowed, low risk, no approval neededMCP call + JWT + tenant contextGET /sim/v1/services/{id}/healthhealth payloadtyped tool resulttool result appended to contextfinal answer citing the toolanswer

And the write path, which is where the series spends its paranoia:

Lmcp-operations-serverPolicy + approval storeagent-apiOps engineerLmcp-operations-serverPolicy + approval storeagent-apiOps engineeropen an incident for payment-gatewaymodel proposes create_incidentproposed call + argumentsscope check, risk class, approval requiredpending approval + single-use tokenSSE approval request (token, arguments, expiry)POST /approvals/{id} approvetoken valid, single-use, args hash matchescreate_incident + JWT + idempotency keyscope + tenant + schema validationcreated incidentaudit event recordedanswer with incident reference
Lmcp-operations-serverPolicy + approval storeagent-apiOps engineerLmcp-operations-serverPolicy + approval storeagent-apiOps engineeropen an incident for payment-gatewaymodel proposes create_incidentproposed call + argumentsscope check, risk class, approval requiredpending approval + single-use tokenSSE approval request (token, arguments, expiry)POST /approvals/{id} approvetoken valid, single-use, args hash matchescreate_incident + JWT + idempotency keyscope + tenant + schema validationcreated incidentaudit event recordedanswer with incident reference

Security note: the model proposes every tool call, but the decision is never the model’s. Scope check, risk classification, approval requirement, and argument integrity all happen in deterministic code. If the model changes the arguments after approval — or a prompt injection asks it to — the argument-hash check fails and the call is rejected.

Trust boundaries

Untrusted / probabilistic

Trust boundary 2: service-to-service

Trust boundary 1: authenticated callers

Untrusted

JWT over HTTPS

JWT, audience=mcp

service JWT

untrusted text in, untrusted text out

untrusted content, delimited

Client

agent-api

mcp-operations-server

operations-simulator

Model endpoint

Runbook content

Untrusted / probabilistic

Trust boundary 2: service-to-service

Trust boundary 1: authenticated callers

Untrusted

JWT over HTTPS

JWT, audience=mcp

service JWT

untrusted text in, untrusted text out

untrusted content, delimited

Client

agent-api

mcp-operations-server

operations-simulator

Model endpoint

Runbook content

Prompts, model output, retrieved runbook text, and MCP metadata are all treated as untrusted input. The delimiters in the system prompt are a parsing aid, not a security boundary — the security boundary is the deterministic policy code.

Ingestion and retrieval

agent-api (request path)

opsdb.knowledge

knowledge-ingestion (batch)

read runbooks

semantic chunking

embed

upsert chunks

documents

document_chunks + embeddings

embed question

filtered similarity search

build cited context

agent-api (request path)

opsdb.knowledge

knowledge-ingestion (batch)

read runbooks

semantic chunking

embed

upsert chunks

documents

document_chunks + embeddings

embed question

filtered similarity search

build cited context

Ingestion is a separate command-line module, not a side effect of an HTTP request: document parsing, section-aware chunking, deterministic chunk IDs, per-document checksums so only changed runbooks are re-embedded.

The advisor chain and agent loop

tool proposal

allowed

bounded loop, max N

needs approval

denied

HTTP request

auth + tenant context

conversation memory

retrieval advisor

tool-calling advisor

structured-output validation

observability advisor

ChatModel

tool policy registry

execute via MCP client

emit approval event, suspend

safe refusal

tool proposal

allowed

bounded loop, max N

needs approval

denied

HTTP request

auth + tenant context

conversation memory

retrieval advisor

tool-calling advisor

structured-output validation

observability advisor

ChatModel

tool policy registry

execute via MCP client

emit approval event, suspend

safe refusal

Every loop iteration consumes budget — model calls, tool calls, wall-clock time, tokens. The loop terminates on an answer, a refusal, an approval wait, or budget exhaustion; there is no unbounded “the model decides when to stop.”

Deployment

Compose is the runnable reference environment:

Docker Compose network

agent-api :8080

mcp-operations-server :8081

operations-simulator :8082

postgres+pgvector :5432

knowledge-ingestion (job)

keycloak :8085

ollama :11434

otel-collector :4317/4318

prometheus :9090

tempo :3200

loki :3100

grafana :3000

Docker Compose network

agent-api :8080

mcp-operations-server :8081

operations-simulator :8082

postgres+pgvector :5432

knowledge-ingestion (job)

keycloak :8085

ollama :11434

otel-collector :4317/4318

prometheus :9090

tempo :3200

loki :3100

grafana :3000

Kubernetes (Chapter 15) mirrors the same shapes with Deployments, a one-shot ingestion Job, ConfigMaps, Secrets, NetworkPolicies, and PodDisruptionBudgets — the Compose file is the honest local environment, not a mini-production.

Threat model skeleton

docs/threat-model/ grows through the series; the skeleton committed at this checkpoint names assets, actors, entry points, and the threats that shape the design.

Assets: tenant-scoped runbooks; live operational data (service health, incidents); the ability to create incidents and notes; credentials and tokens; model spend; audit integrity.

Actors: ops viewer, ops operator, service accounts, a compromised or malicious document author, a prompt-injection attacker, a compromised MCP client.

Threats with their primary controls:

ThreatControl
Direct prompt injectionPrompts never decide authorization; policy code does
Indirect injection via runbook contentRetrieved text delimited and treated as data; ingestion review gate
Cross-tenant retrieval or tool accessTenant filter in retrieval query and tenant policy on every tool call
Confused deputy (agent acts with ambient privileges)Per-tool scopes; service account cannot approve its own writes
Approval replay / TOCTOUSingle-use tokens bound to user, tenant, tool, normalized args hash, expiry
Duplicate incident on retryIdempotency-Key required on all writes, enforced by simulator
Token theft / wrong issuer / wrong audienceJWT issuer + audience + scope validation on both HTTP boundaries
SSRF via model-generated URLs/identifiersTyped identifiers and allowlists; model never supplies raw URLs
Secrets/prompts in logs, traces, evalsRedaction filters; audit sink separate from diagnostics
DoS via long prompts, recursive calls, retry stormsSize limits, step budgets, deadlines, retry classification, bulkheads
Supply chain (deps, models)Locked versions, pinned image tags, pinned evaluator prompts

Residual risks (accepted and documented): a clever injection may still produce a wrong answer — which is why writes are gated but answers are advisory; local Ollama profiles relax some controls for developer convenience, and that relaxation is labeled, not hidden.

Technology choices and alternatives

DecisionChoiceAlternative considered
RuntimeJava 25 LTSJava 27 — Boot 4.1 supports through 26; 27 appears only in an optional lab
FrameworkSpring Boot 4.1.1 / Framework 7.xBoot 3.5 — the 4.x/AI-2.0 line is the current pairing
AI integrationSpring AI 2.0.1 (spring-ai-bom)LangChain4j — fine library; Spring AI keeps the BOM-managed stack coherent
MCP transportStreamable HTTP, stateless toolsSSE transport — deprecated in MCP SDK 2.0; stdio — no network auth boundary
Tool server languageKotlin 2.3.21Java — Kotlin earns its place here (data classes, when); one module, not the whole repo
Vector storePostgreSQL 17 + pgvector 0.8.6Qdrant/pgvector-managed — extra infra for runbook-scale data
IdentityKeycloak 26.7.x, OIDCHand-rolled JWT — never in a security chapter
ResilienceFramework 7 retry/concurrency primitivesResilience4j — added only if Chapter 11 hits its limits
Load testGatlingk6 — JVM DSL keeps everything in one toolchain

Expected resource requirements

A laptop with 8 GB free RAM runs the whole Compose stack: Ollama with qwen3:4b plus nomic-embed-text fits in roughly 4 GB; a llama3.2:3b path is documented for tighter machines. Ports used: 3000, 3100, 3200, 4317/4318, 5432, 8080–8082, 8085, 9090, 11434. The hosted-provider profile is optional and excluded from default builds; no paid account is required anywhere in the series.

Repository roadmap and the final demo

The tree from the blueprint is committed empty in Chapter 1 and filled in chapter order: simulator before agent (Chapter 2 before 3), knowledge before tools (4–5 before 7–8), security before writes (9 before 10). Each chapter ends at a Git tag.

The final exercise in Chapter 16 replays one scenario end to end: a user asks why payment-gateway is degraded; the agent cites the runbook section, calls get_service_status and list_recent_incidents, produces a schema-valid IncidentAssessment, proposes create_incident, waits for human approval, executes exactly once under an idempotency key, and the whole path is visible as one trace in Tempo with an audit record for the write.

What comes next

Chapter 1 turns this map into a buildable skeleton: the Gradle wrapper, version catalog, convention plugins, and the first ArchUnit rules — so that domain-contracts physically cannot drift toward Spring AI no matter what a later chapter adds.

Project State Ledger — chapter-00-architecture

  • Modules declared: build-logic, domain-contracts, agent-api, knowledge-ingestion, mcp-operations-server, operations-simulator, evaluation-suite, architecture-tests, test-support
  • Root package: in.o612.eng.opsagent.<module>; Kotlin escapes in with backticks
  • Ports: agent-api 8080, mcp 8081, simulator 8082, keycloak 8085, postgres 5432, ollama 11434, otel-collector 4317/4318, prometheus 9090, tempo 3200, loki 3100, grafana 3000
  • Databases: simdb (schema sim), opsdb (schemas knowledge, agent)
  • Scopes: agent:invoke; ops:read; ops:incident:write; ops:note:write
  • MCP tools (planned): get_service_status, list_recent_incidents, get_incident, create_incident, append_incident_note
  • Models: Ollama qwen3:4b chat, nomic-embed-text embeddings (vector(768)); hosted profile via spring-ai-starter-model-openai
  • Tenants: acme, globex; users priya (operator), sam (viewer)
  • Next: chapter-01-build-foundation — Gradle scaffold, convention plugins, CI, first architecture tests
Spring BootKotlinJavaAI

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind