Containerize and deploy: Compose for real, Kubernetes without lies
Checkpoint tag: chapter-15-deployment — the whole platform starts with one Compose command; Kubernetes manifests pass kubectl apply --dry-run=server and run on a local cluster.
What will be built
Layered container images for all four services (jlink-free, but honest about the JVM footprint), the complete Compose stack as the reference environment, and deployment/ manifests: base resources plus local and production-example Kustomize overlays — Deployments, Services, an ingestion Job, ConfigMaps, Secret templates (never real secrets), NetworkPolicies, PodDisruptionBudgets, and resource requests/limits sized from Chapter 14’s measurements.
Why it matters
Deployment is where hand-waving becomes YAML. The Compose stack is the environment every earlier chapter’s commands assume; the Kubernetes manifests are where the security model gets infrastructure teeth — the simulator’s fault admin endpoint exists on the network somewhere, and NetworkPolicy is how you guarantee the agent’s side of the mesh can’t reach it. “It worked on Compose” earns nothing until probes, policies, and rollback exist.
Concepts explained
Probes are a contract with the orchestrator. /actuator/health/liveness answers “is the process dead” (restart); /readiness answers “can it serve” (remove from Service endpoints); the ingestion Job needs neither — it’s a batch/v1 Job with restartPolicy: Never, not a Deployment pretending to be one. Wrong probe choices cause the classic “rolling update kills healthy pods during a slow dependency’s warmup.”
Secrets: placeholders with structure. deployment/base/*-secret.yaml files carry stringData keys with REPLACE_ME values and a comment pointing at the real mechanism (External Secrets, sealed-secrets, SOPS — your org’s choice). The repo rule is absolute: no real secret has ever existed in this Git history.
NetworkPolicy mirrors the threat model. agent-api may egress to MCP, Postgres, Keycloak, Ollama/hosted, OTel. mcp-operations-server may egress to the simulator and Keycloak only. The simulator accepts ingress only from MCP. The /sim/admin/** namespace of risk is unreachable from agent-api by policy, not by luck.
Files added or changed
Dockerfile.jvm (shared, parameterized)infra/compose/docker-compose.yml (final full stack)deployment/base/{agent-api,mcp-operations-server,operations-simulator,postgres,keycloak,observability}/*.yamldeployment/base/jobs/knowledge-ingestion-job.yamldeployment/base/{network-policies.yaml, pdb.yaml, secrets.example.yaml}deployment/overlays/{local,production-example}/kustomization.yamlscripts/{deploy-local.sh, smoke.sh}Complete code (the load-bearing parts)
# Build once per module: docker build --build-arg MODULE=agent-api -f Dockerfile.jvm .FROM eclipse-temurin:25-jre AS runARG MODULEWORKDIR /appCOPY ${MODULE}/build/libs/${MODULE}.jar app.jarUSER 1000:1000EXPOSE 8080ENTRYPOINT ["java","-XX:MaxRAMPercentage=75","-XX:+UseZGC","-jar","app.jar"]apiVersion: apps/v1kind: Deploymentmetadata: { name: agent-api }spec: replicas: 2 strategy: { rollingUpdate: { maxUnavailable: 0, maxSurge: 1 } } template: spec: containers: - name: agent-api image: ops/agent-api:1.0.0 envFrom: [{ configMapRef: { name: agent-api-config } }] env: - name: AGENT_MCP_CLIENT_SECRET valueFrom: { secretKeyRef: { name: agent-api-secrets, key: mcp-client-secret } } resources: requests: { cpu: 250m, memory: 768Mi } limits: { cpu: "1", memory: 1Gi } readinessProbe: httpGet: { path: /actuator/health/readiness, port: 8080 } periodSeconds: 5 livenessProbe: httpGet: { path: /actuator/health/liveness, port: 8080 } periodSeconds: 15 failureThreshold: 3 startupProbe: httpGet: { path: /actuator/health/liveness, port: 8080 } periodSeconds: 5 failureThreshold: 24 # ~2min for JVM + Flyway + model warmupapiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: { name: simulator-ingress }spec: podSelector: { matchLabels: { app: operations-simulator } } policyTypes: [Ingress] ingress: - from: - podSelector: { matchLabels: { app: mcp-operations-server } } ports: [{ port: 8082 }]The ingestion Job runs knowledge-ingestion with INGEST_DRY_RUN=false on a ttlSecondsAfterFinished cleanup, mountable runbook ConfigMap for the demo corpus — in production, runbooks arrive via CI-published artifact, which the docs note as the real-world hook.
Compose: the full stack
docker-compose.yml now carries postgres, ollama, keycloak (realm import mounted), operations-simulator, mcp-operations-server, agent-api, otel-collector, prometheus, tempo, loki, grafana — with depends_on: { condition: service_healthy } chains so up is genuinely one command. The agent-api service gets MODEL_PROVIDER=ollama and internal DNS names (http://mcp-operations-server:8081); per-hostname properties live in a .env.example file, real env exported by the operator.
Rollout, rollback, and the honest limits
Rolling update with maxUnavailable: 0; rollback is kubectl rollout undo because nothing here carries schema risk that isn’t Flyway-gated — and Flyway forward-only migrations mean “roll back code, forward-fix schema” is the stated policy. Autoscaling: HPA on CPU is included but annotated — model-call latency dominates so replicas scale on concurrency metrics (agent.model.bulkhead saturation) once you have real traffic; CPU-based HPA is the placeholder, labeled as such.
Known limitations, printed rather than buried: Compose has no HA story (it’s a dev environment); K8s overlay assumes a storage class and ingress controller exist; Ollama in-cluster wants a GPU node pool — the production-example overlay points the model endpoint at an external host for that reason.
Failure-injection lab
kubectl delete podonmcp-operations-servermid-request → agent returnsUPSTREAM_UNAVAILABLEresults; readiness gates the new pod before it serves.- Scale simulator to 0 → read tools degrade; the agent still answers runbook questions — the blast radius is the tool namespace, not the service.
- Apply the NetworkPolicy → confirm
/sim/admin/faults/outageis unreachable fromagent-api(kubectl exec+ curl), still reachable from within the simulator pod.
Checkpoint verification checklist
-
docker compose upreaches all-healthy without manual ordering. -
kubectl apply -k overlays/localconverges on kind/minikube. - Probes distinguish liveness from readiness; Job ≠ Deployment.
- NetworkPolicy blocks agent→admin path; zero secrets in Git.
Commit message and Git tag
feat(deploy): layered images, full compose stack, kustomize base + overlaysgit tag chapter-15-deployment
What comes next
Chapter 16 is the review: the threat model revisited against the built system, the operational drills, and the final incident scenario that exercises everything at once.
Project State Ledger — chapter-15-deployment
- Images: shared
Dockerfile.jvm, Temurin 25 JRE, ZGC, MaxRAMPercentage=75, non-root uid 1000 - Compose: full 11-service stack, health-gated ordering
- K8s: base + local/production-example overlays; ingestion as Job; NetworkPolicies enforce the trust boundaries; no real secrets in repo
- SLO/HPA: CPU-HPA placeholder flagged; saturation-based scaling documented as the real signal
- Next:
chapter-16-hardening→v1.0.0