Series overview
Part 17 of 2861% complete
2026-06-30•16 min read

Cache-aside and distributed caching

Chapter 16 gave order-service’s calls to inventory-service timeouts, retries, circuit breaking, and bulkheads — protection against inventory-service being slow or down. This chapter addresses a different problem: inventory-service being perfectly healthy but still buckling under the sheer volume of reads for the same few dozen SKUs during a flash sale.

1. Problem the Pattern Solves

During a flash sale on a limited-quantity item, inventory-service’s database receives tens of thousands of available(sku) reads per second for the same handful of SKUs — the product page’s “12 left in stock” display alone accounts for most of it, refreshed on every page load. The database isn’t down and isn’t even particularly slow per query, but the sheer read volume against a handful of hot rows creates lock contention that measurably increases latency for the (comparatively rare) actual stock-decrementing writes — the operations that matter most for correctness.

The straightforward fix — cache the stock level — has an obvious trap: cache it too eagerly, or for too long, and the product page shows “12 left in stock” to a hundred customers after the real count has dropped to zero, leading to accepted orders that inventory-service’s actual reservation logic (Chapter 1’s reserveStock, still the single source of truth) will reject moments later — a worse customer experience than no caching at all, in one specific direction (over-promising availability).

Forces in tension:

  • Read performance vs. staleness risk. A cache dramatically reduces database read load, but every cached value is, by definition, a snapshot that can be stale by the time it’s read — the core trade-off this pattern can never fully eliminate, only manage.
  • Cache hit rate vs. correctness for the operation that actually matters. The display-only stock count can tolerate a few seconds of staleness (a customer sees “12 left” when it’s really 9); the actual reservation decision inside reserveStock absolutely cannot use a stale cached value, or it would double-reserve stock that’s already gone.
  • Operational simplicity vs. a new failure mode. Redis (or any cache) is a new piece of infrastructure with its own availability profile — a cache outage must degrade to “slower, but correct” (reading the database directly), never to “fast, but wrong.”
  • Cache invalidation complexity. Every write path that changes a cached value now has to remember to invalidate or update the cache — a coordination burden that grows with the number of writers and is a notorious source of subtle bugs if any writer forgets.

2. Core Idea

Cache-aside (also called lazy loading) means the application code, not the cache itself, is responsible for reading from and populating the cache: on a read, check the cache first; on a miss, read from the database and populate the cache for next time; on a write, invalidate (or update) the corresponding cache entry so subsequent reads don’t see stale data.

Write path

reserveStock() decrements DB

DEL stock:WIDGET-1

(never trust a stale write-through)

Read path (cache-aside)

yes

no

GET stock:WIDGET-1

in cache?

return cached value

query inventory_schema

SET stock:WIDGET-1

(TTL: 2s)

return value

Write path

reserveStock() decrements DB

DEL stock:WIDGET-1

(never trust a stale write-through)

Read path (cache-aside)

yes

no

GET stock:WIDGET-1

in cache?

return cached value

query inventory_schema

SET stock:WIDGET-1

(TTL: 2s)

return value

Participants:

  • Cache — Redis, here, chosen for its low latency, wide Spring support, and because Northwind already runs it for the API gateway’s rate limiter (Chapter 4) — reusing existing operational knowledge rather than introducing a second caching technology.
  • Cache-aside logic — application code deciding when to read from cache versus database, and when to invalidate — deliberately not pushed into the cache or database themselves, keeping the caching decision visible and testable in inventory-service’s own code.
  • Source of truth — inventory_schema (Chapter 1, unchanged) remains authoritative; the cache is always a derived, disposable, rebuildable copy.

Commonly confused with:

  • Write-through / write-behind caching. These alternatives update the cache as part of the write path (write-through: synchronously; write-behind: asynchronously), keeping the cache always populated. Cache-aside instead invalidates on write and repopulates lazily on the next read — simpler to reason about and less prone to cache/database divergence from a missed write-through call, at the cost of every post-invalidation read paying one cache-miss database query. This chapter uses cache-aside because inventory-service’s read pattern (many reads, comparatively few writes, tolerant of a short TTL) fits it well; a workload with the opposite ratio might reasonably choose differently.
  • CQRS’s read projections (Chapter 10). A CQRS projection is a durable, purpose-built read model kept in sync with a write model, often expected to be complete and long-lived. A cache-aside entry is disposable, short-lived (Section 5’s 2-second TTL), and exists purely for performance — losing every cache entry right now would be a performance blip, never a correctness or data-loss problem, unlike losing a CQRS projection’s data.
  • Session storage. Using Redis (or any cache) to store user session state is a different, stateful use case with different consistency requirements (a lost session actually loses something) than cache-aside’s disposable, rebuildable performance optimization.

3. When to Use It

Strong indicators:

  • A small number of frequently-read values (hot keys) account for a disproportionate share of read load — Northwind’s flash-sale SKUs are the textbook case.
  • The cached data can tolerate a short, bounded staleness window for its read use case, even if the write (authoritative) path cannot.
  • Read volume is measurably causing database contention or latency that a cache would relieve — evidence, not assumption, exactly the standard this series has applied to every pattern so far.

Concrete use cases:

  • E-commerce, as here: product availability display, pricing, and catalog data are classic cache-aside candidates — high read volume, tolerant of brief staleness for display purposes.
  • Content platforms: article or media metadata, viewed far more often than it’s edited, caches well with cache-aside and a moderate TTL.
  • SaaS configuration/feature-flag data: read on nearly every request, changed rarely — an excellent fit, provided the flag-evaluation logic can tolerate the cache’s TTL-bounded staleness (or uses an explicit invalidation push instead of relying purely on TTL expiry).
  • Any read replica-adjacent workload where the data is simple key-value shaped and doesn’t need the relational query flexibility a database read replica would provide.

Prerequisites:

  • A clear, explicit answer to “which reads may use the cache and which must go to the source of truth” — Northwind’s product-page display versus reserveStock’s internal decision is exactly this split, and getting it wrong in the wrong direction (using cached data for the reservation decision) is a correctness bug, not just a performance one.
  • A cache technology already operable at production standards, or budget to bring one to that standard — Redis, reused here from Chapter 4’s rate limiter, rather than introducing yet another piece of infrastructure for a second, unrelated need.
  • An invalidation strategy for every write path that changes cached data — a single missed invalidation is a silent, hard-to-detect correctness bug (Section 8).

4. When Not to Use It

  • The reservation/write decision itself. reserveStock’s core logic — checking current stock and decrementing it — must always read the authoritative database value, inside its own transaction, never a cached one. Using a cached stock level for this decision is exactly the “phantom in-stock” bug Section 1 describes, and no TTL is short enough to make it safe, since even a one-second-stale read can double-sell the last unit during a burst.
  • Low read volume, no measured database contention. Caching a rarely-read value adds invalidation complexity and a new failure mode for no performance benefit — evaluate against actual measured load, as with every pattern in this series.
  • Data that changes on every read, or whose staleness tolerance is effectively zero. A live auction’s current highest bid, for instance, has no meaningful caching window — the “freshest” value is the only correct one, making cache-aside actively harmful there.
  • Overengineering signal: introducing a distributed cache for a single-instance service with no measured latency or database-load problem, purely because “production systems use caching.” Match this pattern to a demonstrated need, exactly as the rest of this series has insisted.

5. Implementation Example

The cache-aside read path, used only for the product-page display, never for the reservation decision:

inventory-service/build.gradle.kts
dependencies {
implementation("org.springframework.boot:spring-boot-starter-data-redis")
}
inventory-service/src/main/kotlin/in/o612/eng/northwind/inventory/internal/StockDisplayService.kt
package `in`.o612.eng.northwind.inventory.internal
import org.springframework.data.redis.core.StringRedisTemplate
import org.springframework.stereotype.Service
import java.time.Duration
/** Serves display-only stock counts (product pages, search results).
* Never used by reserveStock() — see InventoryService below, which reads
* the database directly and is the sole source of truth for reservations. */
@Service
class StockDisplayService(
private val redis: StringRedisTemplate,
private val stockRepository: StockRepository,
) {
private val ttl = Duration.ofSeconds(2) // short enough that "12 left" is rarely wrong by more than a couple of units
fun displayedStockLevel(sku: String): Int {
val cacheKey = "stock-display:$sku"
redis.opsForValue().get(cacheKey)?.let { return it.toInt() }
val current = stockRepository.available(sku)
redis.opsForValue().set(cacheKey, current.toString(), ttl)
return current
}
}

The write path invalidates, never write-throughs a possibly-wrong value — deliberately simple, favoring correctness (a cache miss and one extra database read) over the complexity of trying to compute and write the exact new cached value from inside a concurrent write path:

inventory-service/src/main/kotlin/in/o612/eng/northwind/inventory/internal/InventoryService.kt (revised from Chapter 1)
package `in`.o612.eng.northwind.inventory.internal
import org.springframework.data.redis.core.StringRedisTemplate
import org.springframework.stereotype.Service
import org.springframework.transaction.annotation.Transactional
@Service
internal class InventoryService(
private val stockRepository: StockRepository,
private val redis: StringRedisTemplate,
// ... event publisher, unchanged from Chapters 1, 7, 12
) {
@Transactional
fun reserveStock(orderId: java.util.UUID, items: List<StockReservationRequest>): ReservationResult {
// Reads stockRepository directly — the authoritative source —
// never the display cache. This is the one read path in the
// whole service that must never be cache-aside.
val shortages = items.filter { stockRepository.available(it.sku) < it.quantity }
if (shortages.isNotEmpty()) return ReservationResult.Unavailable(shortages.map { it.sku })
items.forEach { stockRepository.decrement(it.sku, it.quantity) }
items.forEach { redis.delete("stock-display:${it.sku}") } // invalidate, don't write-through
return ReservationResult.Reserved
}
}

Invalidating (DEL) rather than write-through (SET with the new value) is a deliberate simplicity choice: computing the exact post-decrement value to write through correctly under concurrent decrements is a harder problem than just letting the next read pay one cache-miss database query — and Northwind’s read volume is high enough that the resulting cache-miss rate after an invalidation is a negligible fraction of total traffic.

Graceful degradation when Redis itself is unavailable — the cache failing must never mean the feature fails, only that it gets slower:

inventory-service/src/main/kotlin/in/o612/eng/northwind/inventory/internal/StockDisplayService.kt (resilient version)
package `in`.o612.eng.northwind.inventory.internal
import org.springframework.data.redis.core.StringRedisTemplate
import org.springframework.stereotype.Service
import org.slf4j.LoggerFactory
import java.time.Duration
@Service
class StockDisplayService(private val redis: StringRedisTemplate, private val stockRepository: StockRepository) {
private val log = LoggerFactory.getLogger(javaClass)
private val ttl = Duration.ofSeconds(2)
fun displayedStockLevel(sku: String): Int {
val cacheKey = "stock-display:$sku"
val cached = runCatching { redis.opsForValue().get(cacheKey) }
.onFailure { log.warn("Redis unavailable, falling back to direct DB read for {}", sku, it) }
.getOrNull()
if (cached != null) return cached.toInt()
val current = stockRepository.available(sku)
runCatching { redis.opsForValue().set(cacheKey, current.toString(), ttl) }
.onFailure { log.warn("Redis unavailable, skipping cache population for {}", sku, it) }
return current
}
}

A Redis outage now degrades StockDisplayService to hitting the database directly on every call — slower, exactly as it was before this chapter, but never wrong and never failing the request.

6. Step-by-Step Flow

reserveStock (checkout)inventory_schemaRedisStockDisplayServiceCustomer (product page)reserveStock (checkout)inventory_schemaRedisStockDisplayServiceCustomer (product page)Meanwhile, a real checkout happensview product page (WIDGET-1)GET stock-display:WIDGET-1missSELECT available FROM stock WHERE sku='WIDGET-1'12SET stock-display:WIDGET-1 = 12 (TTL 2s)"12 left in stock"reserveStock reads DB directly (never the cache)available = 12, decrement to 11DEL stock-display:WIDGET-1page auto-refreshesGET stock-display:WIDGET-1miss (invalidated)SELECT available11"11 left in stock" — correct, one read cycle later
reserveStock (checkout)inventory_schemaRedisStockDisplayServiceCustomer (product page)reserveStock (checkout)inventory_schemaRedisStockDisplayServiceCustomer (product page)Meanwhile, a real checkout happensview product page (WIDGET-1)GET stock-display:WIDGET-1missSELECT available FROM stock WHERE sku='WIDGET-1'12SET stock-display:WIDGET-1 = 12 (TTL 2s)"12 left in stock"reserveStock reads DB directly (never the cache)available = 12, decrement to 11DEL stock-display:WIDGET-1page auto-refreshesGET stock-display:WIDGET-1miss (invalidated)SELECT available11"11 left in stock" — correct, one read cycle later
  1. Client action. A customer views the product page; a separate customer, concurrently, completes a checkout for the same SKU.
  2. API request. The product-page view hits StockDisplayService, which checks Redis first.
  3. Service behavior. On a cache miss, StockDisplayService reads the database and populates the cache with a short TTL.
  4. Database interaction. The checkout’s reserveStock call reads and writes the database directly, entirely bypassing the display cache — the authoritative path this pattern never touches.
  5. Inter-service communication. None new here — this pattern operates entirely within inventory-service’s own boundary.
  6. Error or failure handling. If Redis is unreachable, StockDisplayService falls back to a direct database read (Section 5’s resilient version) — degraded latency, never degraded correctness.
  7. Observability signals. Track cache hit rate and invalidation frequency per key pattern — a hit rate near zero suggests the TTL is too short or the key is genuinely not hot enough to benefit from caching at all.
  8. Final response. The product page briefly shows a slightly stale count (a normal, accepted trade-off for display), while the actual reservation decision — the operation that matters for correctness — was never at risk of using stale data at any point in this flow.

7. Production Concerns

  • Timeouts, retries, idempotency. Cache reads/writes should have their own short timeout, distinct from the database’s — a slow or partitioned Redis instance shouldn’t make a cache-aside read slower than skipping the cache entirely would have been; fail fast to the database fallback (Section 5) rather than waiting on a struggling cache.
  • Data consistency and transaction boundaries. The invalidation (DEL) must happen after the database write commits, not before or concurrently — an invalidation that races ahead of the actual commit could let a stale read repopulate the cache with the old value immediately afterward, a classic cache-aside race condition worth testing for explicitly under concurrent load.
  • TTL selection. Northwind’s 2-second TTL for high-value flash-sale SKUs is a deliberate, narrow choice — a general product catalog with far lower write frequency might reasonably use a much longer TTL (minutes), trading a longer staleness window for a much higher cache hit rate; there’s no universal correct TTL, only one calibrated to each specific value’s write frequency and staleness tolerance.
  • API versioning. Not directly affected — the cache is purely an internal performance detail of inventory-service, invisible to any external contract.
  • Authentication and service-to-service trust. The Redis instance should be network-isolated to the services that need it, with authentication enabled (Redis AUTH or equivalent) — a shared cache with no access control is a soft target for any compromised service on the same network to read or poison.
  • Logging, metrics, tracing. Log and alert on a sudden drop in cache hit rate or a spike in Redis errors — both are leading indicators of either a cache-warming problem after a deploy or a genuine Redis-side incident.
  • Kubernetes deployment, health probes, autoscaling. Redis should run as its own managed or self-hosted deployment with its own health checks and, ideally, replication for availability — inventory-service’s own readiness probe should not depend on Redis being reachable, precisely because Section 5’s graceful-degradation design means the service functions correctly without it, just slower.
  • Testing strategy. Test the cache-aside logic against an embedded or Testcontainers-managed Redis instance, explicitly including: a cache-miss path, a cache-hit path, an invalidation-then-re-read path, and a Redis-unavailable path verifying graceful degradation — each a distinct, necessary test case.
  • Migration strategy. Introduce caching for the single highest-read-volume key pattern first (Northwind’s flash-sale stock display), verify its hit rate and correctness under real load, and expand to other read-heavy paths (catalog data, pricing) only once each is independently justified by measured load — not as a blanket “cache everything” rollout.

8. Common Mistakes

  1. Using cached data for a write/reservation decision. Reading StockDisplayService’s cached value inside reserveStock, even “just as a quick pre-check before the real database read,” risks a stale value influencing a decision that must be authoritative. Fix: the reservation path never touches the cache at all, as Section 5’s InventoryService.reserveStock demonstrates.
  2. Write-through logic that races with concurrent writes. Attempting to compute and cache the new stock value directly inside the write path, under concurrent decrements, risks caching a value that’s already wrong by the time the SET completes. Fix: invalidate (DEL) rather than write-through, as Section 5 does, accepting one extra cache-miss query in exchange for correctness.
  3. No graceful degradation when the cache is unavailable. Letting a Redis connection failure propagate as an exception that fails the product-page request treats a performance optimization’s failure as if it were a correctness-critical dependency’s failure. Fix: wrap cache operations to degrade to a direct database read on failure, never to a request failure (Section 5).
  4. TTL chosen without regard to the specific value’s write frequency. Applying one blanket TTL (say, 5 minutes) across both a rarely-changing catalog description and a rapidly-changing flash-sale stock count produces an unacceptably stale flash-sale display. Fix: calibrate TTL per data type’s actual write frequency and staleness tolerance, as Section 7 discusses.
  5. A missed invalidation on one of several write paths. If a second write path is later added (say, a warehouse-reconciliation batch job that also adjusts stock, from Chapter 1’s exercises) and its author doesn’t know to invalidate the display cache, the cache silently diverges from the database with no error anywhere. Fix: centralize cache invalidation inside the repository or service layer itself, so every writer automatically triggers it, rather than relying on every caller to remember.
  6. Caching a value with effectively zero staleness tolerance. Applying cache-aside to data where even a one-second-old value is meaningfully wrong (Section 4’s live-auction example) misapplies a pattern built for tolerant staleness to a case that has none. Fix: verify actual staleness tolerance for the specific data before introducing caching, not just its read volume.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
A small number of hot keys account for disproportionate read load, tolerant of brief stalenessYesCache-aside relieves the database with minimal added complexity—
The read feeds a write/business decision that must be authoritativeNo, for that specific readStale cached data used in a decision is a correctness bug, not a performance trade-offAlways read the source of truth directly for decisions
Low read volume, no measured database contentionNoAdded invalidation complexity with no performance benefitRead the database directly
Data has effectively zero staleness toleranceNoAny caching window is already too stale for the use caseRead the source of truth directly
High read volume and a durable, purpose-built read shape is needed long-term, not just performanceConsider insteadA CQRS projection (Chapter 10) may fit better than a disposable, TTL-based cacheCQRS read model

10. Hands-On Exercise

Extend it: add cache-aside caching for the product catalog’s descriptive data (name, description, image URL) — data that changes far less often than stock levels — and choose and justify an appropriate TTL, contrasting it explicitly with the flash-sale stock display’s 2-second TTL.

Simulate a failure: stop the Redis instance entirely while StockDisplayService is under load, and confirm the service degrades to direct database reads without failing any requests — then verify the cache repopulates correctly and hit rate recovers once Redis comes back.

Decision question, with justification required: Northwind’s marketing team wants the product page’s “X left in stock” display to be exactly accurate at all times during flash sales, arguing the current 2-second staleness window creates a bad customer experience when the last unit sells and the display doesn’t update immediately. Should the TTL be reduced further, should the cache be replaced with a push-based invalidation (an event fired on every decrement, updating the cache immediately rather than just deleting it), or should the display accept its current staleness as a reasonable trade-off? Weigh the options against Section 1’s forces.

11. Key Takeaways

  • Cache-aside means application code explicitly checks the cache, falls back to the database on a miss, and populates the cache — giving full visibility and control over caching decisions, unlike write-through or write-behind alternatives.
  • Never use cached data for a write or business decision that must be authoritative — the reservation logic that actually matters for correctness always reads the source of truth directly, with no exception.
  • Invalidate on write rather than attempting to write-through a computed new value — simpler to reason about correctly under concurrent writes, at the cost of one cache-miss query per invalidation.
  • A cache’s unavailability must degrade performance, never correctness — always design an explicit fallback to the source of truth, never let a cache failure fail the request.
  • Calibrate TTL per data type’s actual write frequency and staleness tolerance — there’s no universal correct value, and applying one blanket TTL across differently-volatile data produces either unnecessary staleness or unnecessary cache-miss overhead.
  • Centralize invalidation logic in the service or repository layer so every write path triggers it automatically — a missed invalidation on one of several writers is a silent, hard-to-detect correctness bug.
  • Introduce caching for a specific, measured hot-read problem — not as a blanket, speculative optimization applied uniformly across a service’s data.
Spring BootKotlinMicroservices

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind