Series overview
Part 24 of 2886% complete
2026-07-12•17 min read

Blue-green, canary, and feature-flag-based deployment

Every deployment in this series so far has been a plain Kubernetes rolling update — new pods gradually replace old ones, all receiving traffic identically. This chapter asks what happens when Northwind needs to deploy a change risky enough that “gradually replace, no differentiation” isn’t good enough.

1. Problem the Pattern Solves

inventory-service’s team has rewritten the core stock-reservation algorithm to fix a subtle race condition under extremely high concurrency (the exact kind of flash-sale load this series has referenced since Chapter 1). The new algorithm passes every unit test, every contract test (Chapter 23), and a thorough load test in staging — but staging’s load test, however thorough, is a simulation, and the team knows from experience that real flash-sale traffic patterns have surprised them before. A plain rolling update would replace 100% of inventory-service’s pods with the new algorithm over a few minutes, meaning if it has a problem only real production load surfaces, every single flash-sale customer is affected before anyone notices.

A related but different need: Northwind’s product team wants to test a new “express checkout” flow (skipping the order-confirmation screen for repeat customers) with a small percentage of real users to measure its effect on conversion, before committing to it for everyone — a decision that has nothing to do with deployment risk and everything to do with product experimentation, but needs a similar mechanism: control precisely who sees the new behavior.

Forces in tension:

  • Deployment risk vs. rollback speed. A rolling update, once fully complete, means rolling back requires another full rolling update in reverse — not instant, and requiring old images to still be available and compatible with any interim data changes.
  • Blast radius vs. deployment complexity. Exposing a change to 100% of traffic immediately is simplest to implement but maximizes the blast radius of an undetected problem; exposing it to a small percentage first requires more sophisticated traffic control but limits exposure while confidence builds.
  • Deployment risk vs. product experimentation — genuinely different problems, often solved by similar mechanisms. Canary deployment answers “is this new code safe” (an engineering risk question); feature flags answer “should this user see this behavior” (a product and business question) — conflating them, as many teams do, can produce a system where nobody’s sure whether a flag exists for safety or for product experimentation.
  • Statistical confidence vs. speed of rollout. A canary needs to run long enough, at enough traffic volume, to actually detect a problem statistically — rolling out to 100% too quickly defeats the purpose; leaving a canary running too long delays the full rollout’s benefit for everyone.

2. Core Idea

Three related but distinct techniques for controlling exposure to a change:

  • Blue-green deployment: run two complete, independent environments (“blue” — currently live, and “green” — the new version), fully tested in isolation, then switch all traffic from blue to green atomically (typically via a load balancer or router change) — instant cutover, and instant rollback by switching back, at the cost of running two full environments simultaneously during the transition.
  • Canary deployment: gradually shift a small, increasing percentage of real traffic to the new version while most traffic still goes to the old one, monitoring the canary’s error rate and latency closely before proceeding to a larger percentage — slower than blue-green’s atomic switch, but detects problems on a small fraction of real traffic before they affect everyone.
  • Feature flags: decouple code deployment from feature activation entirely — the new code ships to production for everyone, but a runtime flag (evaluated per-request, often per-user or per-cohort) controls whether the new behavior actually executes, allowing product-level experimentation and instant, code-deploy-independent toggling.

Feature flag: express checkout

yes

no

order-service (single version, deployed to 100%)

flag: express-checkout

enabled for this user?

skip confirmation screen

standard flow

Canary: inventory-service's new algorithm

95% of traffic

5% of traffic

monitored closely

Load balancer

inventory-service v1 (stable)

inventory-service v2 (canary)

error rate, latency

Feature flag: express checkout

yes

no

order-service (single version, deployed to 100%)

flag: express-checkout

enabled for this user?

skip confirmation screen

standard flow

Canary: inventory-service's new algorithm

95% of traffic

5% of traffic

monitored closely

Load balancer

inventory-service v1 (stable)

inventory-service v2 (canary)

error rate, latency

Participants:

  • Traffic-splitting mechanism — for the canary, this chapter uses the service mesh’s traffic management (Chapter 21’s Istio, already in place) to split a percentage of real traffic by weight; for blue-green, a router or load-balancer switch.
  • Feature flag evaluator — a runtime component (this chapter uses a simple, self-hosted flag service) deciding, per request, whether a flag is active for that specific user or cohort.
  • Monitoring/rollback trigger — the metrics (Chapter 22’s observability infrastructure) and alerting that decide whether a canary proceeds, holds, or rolls back.

Commonly confused with:

  • A/B testing. A/B testing is a product-analytics technique measuring a business outcome (conversion rate) across two variants — it can be implemented using feature flags (as Northwind’s express-checkout experiment does), but the two are not the same thing: a canary deployment’s goal is verifying deployment safety, not measuring a business metric, even though both involve controlled traffic exposure.
  • Rolling updates (used throughout this series since Chapter 1). A plain Kubernetes rolling update replaces old pods with new ones gradually, but every request during the rollout is served by whichever pod happens to receive it — there’s no deliberate, monitored, percentage-controlled exposure decision, and no straightforward way to hold at a specific percentage while evaluating results, unlike a true canary.
  • Chapter 13’s Strangler Fig shadow traffic. Strangler Fig’s shadow traffic mirrors requests to a new implementation without using its response — purely for comparison, with zero risk to real users. A canary does serve real users the new version’s actual response for the canary percentage — genuinely exposing them to it, which is what makes canary analysis meaningful for catching problems shadow traffic’s read-only comparison might miss (e.g., write-path bugs that only manifest when a write actually happens).

3. When to Use It

Strong indicators for canary deployment:

  • A change carries meaningful risk that automated tests (unit, contract, even load tests) may not fully capture — exactly inventory-service’s concurrency-sensitive rewrite, where the team’s own experience says staging simulations have been wrong before.
  • The service already has solid observability (Chapter 22) to actually detect a canary’s problems quickly — a canary without the metrics to evaluate it is just a slower, riskier rolling update.

Strong indicators for blue-green:

  • A change that’s difficult or impossible to canary gradually — a database schema migration that can’t easily run two versions simultaneously, for instance, or a change where partial-traffic exposure would create confusing, inconsistent user experience.
  • A need for a very fast, deterministic rollback (switching a router) rather than a gradual, several-minutes rolling reversal.

Strong indicators for feature flags:

  • A product decision about who sees a feature, independent of when the underlying code was deployed — Northwind’s express-checkout experiment is a clean case: the code is safe to deploy everywhere; whether a specific user experiences the new flow is a business decision, made and changed independently of any deployment.
  • A need to instantly disable a feature in response to a business or operational decision, without waiting for a new deployment.

Concrete use cases:

  • E-commerce, as here: canary a risky algorithm change under real flash-sale load; feature-flag a checkout-flow experiment measured by conversion.
  • Financial services: canary a change to fraud-detection logic against a small percentage of real transactions before trusting it platform-wide, given the cost of both false positives and false negatives at scale.
  • SaaS products: feature-flag a new UI or workflow to a specific customer segment (beta users, a specific plan tier) entirely independent of the underlying code’s deployment schedule.
  • Any platform doing progressive delivery: gradually increasing a canary’s traffic percentage (1% → 10% → 50% → 100%) with automated rollback on metric regression is standard practice for any change where “deploy and hope” carries real business risk.

Prerequisites:

  • Solid observability (Chapter 22) — a canary’s entire value depends on being able to detect its problems quickly and specifically, not just eventually via a customer complaint.
  • A feature-flag evaluation system, self-hosted or third-party, with low-latency evaluation (a flag check shouldn’t meaningfully add to request latency) and a clear ownership/cleanup policy (Section 8) so flags don’t accumulate indefinitely.
  • Traffic-splitting infrastructure — this chapter reuses the service mesh from Chapter 21, since weighted traffic routing is exactly the kind of capability that chapter named as a reason to adopt a mesh at scale.

4. When Not to Use It

  • Low-risk changes with strong automated test coverage and no history of tests missing real problems. A routine dependency bump or a well-tested, low-risk bug fix doesn’t need canary infrastructure’s added process — a plain rolling update, as this series has used throughout, remains appropriate for the majority of changes.
  • Feature flags left in the codebase indefinitely after their experiment concludes. A flag is meant to be temporary — either the feature becomes permanent (remove the flag, keep the code) or it’s rejected (remove the flag and the code) — an accumulating pile of stale, forgotten flags is a maintainability and readability cost with no ongoing benefit (Section 8’s most common mistake).
  • Overengineering signal: canary-deploying every single change, including trivial ones, adds process overhead disproportionate to the actual risk — reserve canary deployment for changes where the team has genuine, articulable uncertainty about production behavior, as Section 3 describes.
  • Risk: running a canary without the statistical rigor to interpret its results correctly — a 1% canary for one hour on a low-traffic service may not generate enough data to detect a real but infrequent problem, giving false confidence before a full rollout.

5. Implementation Example

Canary deployment, using Istio’s traffic-splitting (built directly on Chapter 21’s mesh adoption) to send 5% of inventory-service traffic to the new reservation algorithm:

inventory-service-canary.yaml
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: inventory-service
spec:
hosts: ["inventory-service"]
http:
- route:
- destination: { host: inventory-service, subset: stable }
weight: 95
- destination: { host: inventory-service, subset: canary }
weight: 5
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: inventory-service
spec:
host: inventory-service
subsets:
- name: stable
labels: { version: v1 }
- name: canary
labels: { version: v2 }
Deploying the canary alongside stable
kubectl apply -f inventory-service-v2-deployment.yaml # labeled version: v2, 1-2 replicas only
kubectl apply -f inventory-service-canary.yaml # activates the 95/5 traffic split

Automated canary analysis, comparing the canary’s error rate and latency against the stable baseline using the exact metrics infrastructure Chapter 22 built:

platform-tools/src/main/kotlin/in/o612/eng/northwind/canary/CanaryAnalyzer.kt
package `in`.o612.eng.northwind.canary
import org.springframework.scheduling.annotation.Scheduled
import org.springframework.stereotype.Component
@Component
class CanaryAnalyzer(private val metricsClient: PrometheusClient, private val meshController: MeshTrafficController) {
@Scheduled(fixedDelay = 60_000)
fun evaluateInventoryServiceCanary() {
val stableErrorRate = metricsClient.errorRate(service = "inventory-service", version = "v1")
val canaryErrorRate = metricsClient.errorRate(service = "inventory-service", version = "v2")
val canaryP99 = metricsClient.latencyP99(service = "inventory-service", version = "v2")
val stableP99 = metricsClient.latencyP99(service = "inventory-service", version = "v1")
when {
canaryErrorRate > stableErrorRate * 2 -> {
// Canary is meaningfully worse — automatic rollback, no
// human needs to be paged at 3am to notice this manually.
meshController.rollbackToStable("inventory-service")
alerting.notify("Canary rollback: inventory-service v2 error rate ${canaryErrorRate}, stable ${stableErrorRate}")
}
canaryP99 > stableP99 * 1.5 -> {
meshController.rollbackToStable("inventory-service")
alerting.notify("Canary rollback: inventory-service v2 p99 latency regression")
}
else -> meshController.maybeIncreaseCanaryWeight("inventory-service") // 5% -> 25% -> 50% -> 100%, gradually
}
}
}

This automated gate is what makes canary deployment genuinely safer than a rolling update rather than just a slower one — the rollback decision is made by comparing real metrics against a real, live baseline (the stable version, still receiving 95% of traffic simultaneously), not by hoping someone notices a problem in time.

Feature flags, for the unrelated express-checkout product experiment — deliberately using a different mechanism than the canary above, since it answers a different question:

order-service/build.gradle.kts
dependencies {
implementation("com.flipt:flipt-client-kotlin:1.4.0") // or an equivalent self-hosted flag service
}
order-service/src/main/kotlin/in/o612/eng/northwind/order/api/CheckoutController.kt
package `in`.o612.eng.northwind.order.api
import org.springframework.web.bind.annotation.*
import java.util.UUID
@RestController
class CheckoutController(private val featureFlags: FeatureFlagClient, private val checkoutService: CheckoutService) {
@PostMapping("/api/v1/checkout")
fun checkout(@RequestBody request: CheckoutRequest, @RequestHeader("X-Customer-Id") customerId: UUID): CheckoutResponse {
val expressCheckoutEnabled = featureFlags.isEnabled("express-checkout", context = mapOf("customerId" to customerId.toString()))
return if (expressCheckoutEnabled) {
checkoutService.expressCheckout(request) // skips confirmation screen
} else {
checkoutService.standardCheckout(request) // existing flow, unchanged since Chapter 1
}
}
}

The flag service’s own dashboard controls targeting — say, 10% of returning customers, chosen by the product team, changeable instantly without any deploy — entirely decoupled from order-service’s own release cadence, exactly the property that distinguishes this from the canary mechanism above.

6. Step-by-Step Flow

Canary flow for inventory-service’s reservation algorithm rewrite:

inventory-service v2 (canary, 5%)inventory-service v1 (stable, 95%)CanaryAnalyzerIstio VirtualServicePlatform teaminventory-service v2 (canary, 5%)inventory-service v1 (stable, 95%)CanaryAnalyzerIstio VirtualServicePlatform teamalt[canary healthy][canary regressed]loop[every 60s]Eventually 100% v2, or automatic rollback —either way, decided by real production metrics, not a guessdeploy v2, split 95/5query error rate, p99query error rate, p99increase canary weight (5% -> 25%)rollback to 100% stablealert
inventory-service v2 (canary, 5%)inventory-service v1 (stable, 95%)CanaryAnalyzerIstio VirtualServicePlatform teaminventory-service v2 (canary, 5%)inventory-service v1 (stable, 95%)CanaryAnalyzerIstio VirtualServicePlatform teamalt[canary healthy][canary regressed]loop[every 60s]Eventually 100% v2, or automatic rollback —either way, decided by real production metrics, not a guessdeploy v2, split 95/5query error rate, p99query error rate, p99increase canary weight (5% -> 25%)rollback to 100% stablealert
  1. Client action. A flash-sale customer’s reservation request lands on either v1 or v2 depending on the mesh’s current weight — invisible to the customer either way.
  2. API request. The mesh’s VirtualService routes according to the currently configured split.
  3. Service behavior. Both versions process reservations against the same inventory_schema, so their behavior is directly, fairly comparable under identical real load.
  4. Database interaction. Identical schema and data for both versions — the comparison is purely about the algorithm’s correctness and performance under real concurrency, not any data difference.
  5. Inter-service communication. Unaffected — order-service calls inventory-service exactly as always; it has no awareness a canary is in progress.
  6. Error or failure handling. If v2’s error rate or latency regresses meaningfully against v1’s live baseline, CanaryAnalyzer triggers an automatic rollback within the next evaluation cycle (60 seconds in this configuration) — far faster than a human noticing a dashboard anomaly.
  7. Observability signals. The exact per-version error rate and latency metrics Chapter 22’s instrumentation provides are what make this automated comparison possible at all — this pattern is a direct, concrete payoff of that chapter’s investment.
  8. Final response/outcome. The new algorithm either earns its way to 100% traffic through several rounds of increasing, monitored exposure, or is automatically rolled back with minimal customer impact — a fundamentally different risk profile than a plain rolling update’s all-or-nothing exposure.

7. Production Concerns

  • Timeouts, retries, idempotency. Unaffected directly by the deployment mechanism, though note that Chapter 16’s resilience configuration (retries, circuit breakers) applies identically to both canary and stable versions — a canary’s poor behavior might otherwise be misdiagnosed if its retry configuration accidentally differs from stable’s.
  • Data consistency. Both canary and stable versions write to the same database in this chapter’s example — verify explicitly that a schema or data-shape change isn’t part of the canary (that would need blue-green or a careful migration strategy instead, since two algorithm versions writing incompatible data shapes simultaneously is a real risk).
  • API versioning and backward compatibility. The canary and stable versions must expose identical external contracts to order-service — the canary is testing internal algorithm behavior, not an API change; if the API itself is changing, contract testing (Chapter 23) becomes directly relevant to verify both versions still satisfy consumers.
  • Authentication and service-to-service trust. No new surface — the mesh’s traffic splitting operates below the authentication layer already established in Chapters 4, 6, and 21.
  • Logging, metrics, tracing, correlation IDs. Tag every metric and trace with the serving version (v1/v2) explicitly — this is the specific instrumentation requirement that makes CanaryAnalyzer’s comparison possible; without a version label, stable and canary metrics are indistinguishable.
  • Kubernetes deployment, health probes, autoscaling. The canary deployment needs its own resource allocation and readiness probes, sized appropriately for its (initially small) traffic share — over-provisioning a 5%-traffic canary to the same replica count as the 95%-traffic stable version wastes resources without improving safety.
  • Testing strategy. A canary is not a substitute for pre-deployment testing (unit, contract, load) — it’s an additional, real-production-traffic-based safety net specifically for risks those earlier tests might miss, as Section 1 makes explicit.
  • Migration strategy. Feature flags need an explicit lifecycle: created for an experiment or a risky rollout, evaluated, and then removed — either by making the new behavior permanent (delete the flag and the old code path) or reverting (delete the flag and the new code path) — never left indefinitely in an “it’s just always been there” state.

8. Common Mistakes

  1. Canarying a change with no automated rollback trigger. Deploying a canary and manually watching a dashboard, without an automated comparison like CanaryAnalyzer, means a problem can go unnoticed for far longer than an automated 60-second evaluation cycle would allow. Fix: automate the comparison and rollback decision against the live stable baseline, as Section 5 does.
  2. Confusing a canary’s purpose with a feature flag’s purpose. Using a canary deployment to test a product hypothesis (should we show this UI to some users), rather than to verify code safety, conflates two different questions and typically produces awkward, hard-to-reason-about traffic-splitting logic. Fix: use canaries for deployment-risk questions, feature flags for product/business questions, as Section 2 distinguishes explicitly.
  3. Leaving feature flags in the codebase indefinitely. A flag from an experiment concluded six months ago, still checked on every request, with nobody sure whether it’s safe to remove, is exactly the kind of accumulating technical debt this pattern warns against. Fix: treat every flag’s creation as implicitly scheduling its own removal, tracked with an owner and a decision deadline.
  4. Canarying a change that also alters the database schema or data shape. Running two algorithm versions that write incompatible data simultaneously risks data corruption that no traffic-splitting sophistication can protect against. Fix: use blue-green (an atomic cutover) or a carefully sequenced migration strategy (Chapter 3’s dual-write pattern) for changes that touch data shape, reserving canary deployment for behavior changes that don’t.
  5. Running a canary too briefly or at too small a percentage to reach statistical significance. A 1% canary for five minutes on a service with modest traffic may not generate enough requests to reliably detect an infrequent but real problem, producing false confidence before a full rollout. Fix: calibrate canary duration and traffic percentage to the service’s actual traffic volume and the failure rate you need to be able to detect.
  6. Manually managing traffic-splitting percentages without infrastructure support. Attempting a canary rollout via manual kubectl edits and eyeballed dashboard checks, without the mesh’s declarative traffic management (Chapter 21) or an automated analyzer, is slow, error-prone, and hard to reason about under incident pressure. Fix: build canary deployment on top of infrastructure specifically designed for it (a service mesh’s traffic splitting, as this chapter does), not ad hoc manual steps.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Change carries real, articulable risk automated tests may not fully catchYes (canary)Exposes real production traffic gradually, with automated rollback on regression—
Change involves a schema or data-shape migration that can’t run two versions simultaneouslyYes (blue-green), not canaryAtomic cutover avoids two incompatible data-writing versions running concurrently—
Decision is about which users should see a feature, not about code safetyYes (feature flag)Decouples activation from deployment; a product/business decision, not an engineering-risk one—
Routine, low-risk, well-tested changeNo (any of these)Added process overhead disproportionate to actual riskPlain rolling update
Feature flag’s experiment has concluded, decision madeNo, remove itAn indefinitely-lived flag is pure maintenance cost with no ongoing benefitRemove the flag and the losing code path

10. Hands-On Exercise

Extend it: configure CanaryAnalyzer to also compare a business metric — successful-reservation rate, not just error rate and latency — since a subtly wrong algorithm might return 200 OK responses that are nonetheless functionally incorrect (reserving the wrong quantity, say) without tripping a pure error-rate or latency alert.

Simulate a failure: deploy a deliberately broken inventory-service v2 (returning 500 for 10% of requests) behind the canary configuration, and confirm CanaryAnalyzer detects the elevated error rate and triggers rollback within its configured evaluation window.

Decision question, with justification required: Northwind’s express-checkout feature flag experiment has now run for three months with consistently positive conversion results, and the product team wants to make it permanent for all customers. What specific steps does “making it permanent” actually involve, beyond just setting the flag to 100% — and why does leaving the flag at 100% forever, rather than removing it and the old code path, violate Section 8’s guidance?

11. Key Takeaways

  • Canary deployment, blue-green deployment, and feature flags solve related but distinct problems: canary and blue-green manage deployment risk (is this code safe), while feature flags manage feature activation (who should see this behavior) — decoupled from any specific deployment.
  • A canary’s real value comes from automated comparison against a live stable baseline, using the same production traffic and the same observability infrastructure (Chapter 22) — a canary watched only manually is slower and less reliable than one gated by automated metrics comparison.
  • Reserve blue-green’s atomic cutover for changes (schema migrations, data-shape changes) that can’t safely run two versions simultaneously against the same data — canary deployment assumes both versions can coexist correctly.
  • Every feature flag should have an explicit lifecycle: created for a reason, evaluated, and then removed — either by making the new behavior permanent or reverting it — never left indefinitely as accumulating technical debt.
  • Calibrate a canary’s traffic percentage and duration to the service’s actual traffic volume and the failure rate you need to reliably detect — too small or too brief produces false confidence.
  • Never canary a change that also alters data shape or schema in a way two concurrently-running versions would handle incompatibly — that risk needs blue-green or a dedicated migration strategy instead.
  • These patterns are a direct, concrete payoff of the observability investment from Chapter 22 and the traffic-management capability from Chapter 21’s service mesh — they depend on infrastructure this series built specifically to make deployment risk manageable, not guessed at.
Spring BootKotlinMicroservicesKubernetes

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind