Series overview
Part 19 of 2868% complete
2026-07-03•15 min read

Dead-letter queues and retry topics

Chapter 18 gave notification-service the ability to scale email-sending throughput across competing consumers. It quietly assumed every send eventually succeeds. This chapter addresses what happens when one doesn’t — a malformed email address, a customer’s mail server rejecting the message permanently, or a temporary outage at Northwind’s email provider.

1. Problem the Pattern Solves

A batch of orders comes in from customers who mistyped their email addresses at checkout — a real, unavoidable category of input that will always exist. notification-service’s OrderConfirmationListener (Chapter 18) throws an exception for each one, since the email provider’s API rejects an invalid address outright. With no special handling, Spring Kafka’s default behavior retries the failed message a few times and then — critically — the consumer’s offset doesn’t advance past it, meaning the entire partition is stuck reprocessing the same failing message indefinitely, blocking every other order’s confirmation email queued behind it on that partition, even though those emails have nothing wrong with them.

A second, different failure mode: Northwind’s email provider has a ten-minute outage. Every send fails during that window, for otherwise-perfectly-valid emails — these should eventually succeed once the provider recovers, unlike the malformed-address case, which will never succeed no matter how many times it’s retried.

Forces in tension:

  • Not blocking healthy work vs. not silently dropping failed work. A message that can never succeed must not be allowed to block the many messages behind it that could — but simply discarding it loses a customer’s confirmation email with no record and no recovery path.
  • Distinguishing transient from permanent failure. Retrying a permanently-invalid email address is wasted effort that delays recognizing the real problem; giving up immediately on a transient provider outage abandons work that would have succeeded moments later.
  • Operational visibility vs. automatic handling. Messages that end up unable to be processed need a human-reviewable destination — silently retrying forever, or silently dropping, both hide a real problem (a bad email address that needs correcting, a provider integration that needs fixing) from anyone who could act on it.
  • Cost of the dead-letter/retry infrastructure vs. the failure rate that justifies it. A dead-letter topic and a separate retry topic are more infrastructure and more operational surface — justified once real, non-trivial failure rates make “block the partition” or “silently lose work” unacceptable outcomes.

2. Core Idea

A dead-letter queue (DLQ) (in Kafka terms, a dead-letter topic) is a separate topic where messages that fail processing beyond a defined retry limit are routed, instead of blocking the original topic’s partition or being silently discarded — preserving the failed message for investigation, manual reprocessing, or automated alerting, while letting the main consumer move on to subsequent messages.

A retry topic is a separate topic (or a small chain of topics with increasing delay) used to reattempt a transiently-failed message after a backoff period, without blocking the main topic’s consumer while waiting — the failed message is republished to the retry topic, consumed again after its delay, and either succeeds, is retried further, or eventually lands in the DLQ if retries are exhausted.

success

transient failure

after 30s delay

success

still failing, attempt 2

after 5m delay

success

permanent failure or retries exhausted

order-confirmations

(main topic)

OrderConfirmationListener

email sent

order-confirmations-retry-30s

retry consumer

order-confirmations-retry-5m

retry consumer

order-confirmations-dlq

Alert + manual review dashboard

success

transient failure

after 30s delay

success

still failing, attempt 2

after 5m delay

success

permanent failure or retries exhausted

order-confirmations

(main topic)

OrderConfirmationListener

email sent

order-confirmations-retry-30s

retry consumer

order-confirmations-retry-5m

retry consumer

order-confirmations-dlq

Alert + manual review dashboard

Participants:

  • Main topic consumer — notification-service’s existing OrderConfirmationListener (Chapter 18), now classifying failures rather than treating them uniformly.
  • Retry topic(s) — one or more topics representing increasing backoff stages, each consumed after its configured delay — a small chain (30s, then 5m, in this chapter’s example) rather than an unbounded retry loop.
  • Dead-letter topic — the final destination for messages that exhaust all retry stages or are classified as permanently unrecoverable on the first attempt (a malformed address doesn’t need a 30-second wait to know it will fail again).
  • DLQ consumer/dashboard — a separate, low-volume process or UI that surfaces dead-lettered messages for human review, alerting, or manual reprocessing.

Commonly confused with:

  • Resilience4j’s retry mechanism (Chapter 16). Chapter 16’s retry operates on a synchronous call within a single request’s lifetime (a few attempts, milliseconds to seconds apart, all before the caller gets a response). This chapter’s retry topics operate on asynchronous message processing, with delays measured in seconds to minutes, explicitly decoupled from any caller waiting for a response — a fundamentally different time scale and mechanism, even though both are called “retry.”
  • The transactional outbox / inbox (Chapter 12). The outbox guarantees a message is eventually sent; a DLQ handles what happens when a message is delivered but its processing keeps failing. They protect different points in the pipeline and are frequently both present in the same system, addressing different failure windows.
  • Simply increasing max-attempts on the main consumer. A high retry count on the main topic’s own consumer still blocks that partition for the entire retry duration (with Kafka’s default at-least-once semantics, an uncommitted offset stalls the partition) — routing failed messages to a separate retry topic, as this chapter does, is what actually unblocks the main partition’s other messages during the retry wait.

3. When to Use It

Strong indicators:

  • A message-processing failure rate that’s non-trivial but not catastrophic — occasional invalid input, occasional downstream outages — common enough that manual, ad hoc handling doesn’t scale, rare enough that a full redesign of the workflow isn’t warranted.
  • A real distinction between transient failures (worth retrying) and permanent failures (not worth retrying, worth surfacing immediately) exists in the workload, as it does for Northwind’s email sending.
  • The main topic’s throughput matters enough that one poison message blocking its partition is a real, unacceptable cost — true for notification-service given Chapter 18’s flash-sale scaling work.

Concrete use cases:

  • E-commerce, as here: any email, SMS, or webhook-delivery workload has predictable transient failures (a provider outage) and permanent ones (invalid recipient) that benefit from this exact split.
  • Payment webhook processing: a webhook from a payment processor that fails to process due to a temporary database issue should retry; one that fails due to a malformed or unexpected payload should dead-letter immediately for engineering review, not retry indefinitely.
  • Third-party API integrations generally: any integration point with an external system that has its own, independent uptime and its own category of permanently-invalid requests is a natural fit.
  • ETL and batch pipelines: a single malformed record in a large batch shouldn’t halt the entire pipeline’s progress — dead-lettering the bad record while continuing to process the rest is standard practice.

Prerequisites:

  • A reliable way to distinguish transient from permanent failures for the specific workload (Section 5’s exception-type classification) — without this distinction, either every failure gets the full retry treatment (wasting time on permanent failures) or none does (giving up too early on transient ones).
  • A monitored destination and process for dead-lettered messages — a DLQ nobody looks at is equivalent to silently dropping messages, just with extra storage cost.
  • Idempotent processing (the same discipline required since Chapter 1) — a message may be reprocessed after a retry-topic delay, and the eventual successful processing must be safe even if an earlier attempt partially succeeded before failing.

4. When Not to Use It

  • Failures are rare enough that default retry-and-alert handling is sufficient. A service processing a few dozen messages a day, where an occasional failure is manually noticed and fixed without dedicated DLQ infrastructure, may not need this pattern’s added topics and monitoring yet.
  • Every failure in the workload is permanent, with no transient category. If there’s genuinely no failure mode worth retrying (every failure needs human intervention regardless), a DLQ alone — no retry topic chain — is the right, simpler design; don’t build retry infrastructure for failures that will never succeed on a second attempt.
  • The processing isn’t safely idempotent, and can’t be made so. Retrying a non-idempotent operation risks the exact duplicate-side-effect problem this series has warned against since Chapter 1 — fix the idempotency gap first, or this pattern’s retries become a liability rather than a safety net.
  • Overengineering signal: building a long chain of retry topics with many stages and long backoff periods for a workload where a single short retry, or none at all, would suffice. Match the retry chain’s depth and delay to the actual, observed recovery time of the transient failures you’re handling — Northwind’s 30-second-then-5-minute chain is calibrated to its email provider’s typical outage recovery pattern, not chosen arbitrarily.

5. Implementation Example

Failure classification, distinguishing what’s worth retrying from what isn’t — the single most important decision in this pattern:

notification-service/src/main/kotlin/in/o612/eng/northwind/notification/EmailSendException.kt
package `in`.o612.eng.northwind.notification
sealed class EmailSendException(message: String) : RuntimeException(message) {
/** Provider outage, network blip, rate limit — will likely succeed on retry. */
class TransientProviderFailure(message: String) : EmailSendException(message)
/** Malformed address, provider permanently rejected the recipient —
* will never succeed no matter how many times it's retried. */
class PermanentDeliveryFailure(message: String) : EmailSendException(message)
}

Spring Kafka’s non-blocking retry support, configuring the exact retry-topic chain from Section 2’s diagram:

notification-service/src/main/kotlin/in/o612/eng/northwind/notification/KafkaRetryConfig.kt
package `in`.o612.eng.northwind.notification
import org.springframework.context.annotation.Bean
import org.springframework.context.annotation.Configuration
import org.springframework.kafka.annotation.RetryableTopic
import org.springframework.kafka.retrytopic.TopicSuffixingStrategy
import org.springframework.retry.annotation.Backoff
@Configuration
class KafkaRetryConfig
notification-service/src/main/kotlin/in/o612/eng/northwind/notification/OrderConfirmationListener.kt (revised)
package `in`.o612.eng.northwind.notification
import `in`.o612.eng.northwind.order.api.OrderPlacedNotification
import org.springframework.kafka.annotation.DltHandler
import org.springframework.kafka.annotation.KafkaListener
import org.springframework.kafka.annotation.RetryableTopic
import org.springframework.kafka.retrytopic.TopicSuffixingStrategy
import org.springframework.retry.annotation.Backoff
import org.springframework.stereotype.Component
@Component
class OrderConfirmationListener(
private val emailSender: EmailSender,
private val sentEmailLog: SentEmailLog,
private val deadLetterAlertService: DeadLetterAlertService,
) {
@RetryableTopic(
attempts = "3", // main attempt + 2 retries
backoff = Backoff(delay = 30_000, multiplier = 10.0), // 30s, then 5min
exclude = [EmailSendException.PermanentDeliveryFailure::class], // never retry these — straight to DLQ
topicSuffixingStrategy = TopicSuffixingStrategy.SUFFIX_WITH_INDEX_VALUE,
)
@KafkaListener(topics = ["order-confirmations"], groupId = "notification-service", concurrency = "3")
fun onOrderPlaced(notification: OrderPlacedNotification) {
if (sentEmailLog.alreadySent(notification.orderId)) return
emailSender.sendConfirmation(notification.orderId) // throws EmailSendException subtypes on failure
sentEmailLog.recordSent(notification.orderId)
}
@DltHandler
fun onDeadLetter(notification: OrderPlacedNotification, exception: Exception) {
// Never silently drop — always alert and preserve for manual review.
deadLetterAlertService.recordAndAlert(notification.orderId, exception)
}
}

exclude = [PermanentDeliveryFailure::class] is the mechanism that sends a known-permanent failure straight to the dead-letter topic without wasting time on the 30-second and 5-minute retry stages — the classification from Section 5’s exception hierarchy driving the framework’s routing decision directly, rather than treating every failure identically.

notification-service/src/main/kotlin/in/o612/eng/northwind/notification/DeadLetterAlertService.kt
package `in`.o612.eng.northwind.notification
import org.springframework.jdbc.core.JdbcTemplate
import org.springframework.stereotype.Service
import java.util.UUID
@Service
class DeadLetterAlertService(private val jdbc: JdbcTemplate, private val alerting: AlertingClient) {
fun recordAndAlert(orderId: UUID, exception: Exception) {
jdbc.update(
"INSERT INTO dead_lettered_notifications (order_id, failure_reason, occurred_at) VALUES (?, ?, now())",
orderId, exception.message,
)
// Low-volume, non-urgent alert — a human reviews and decides:
// correct the address and manually resend, or accept the loss.
alerting.notify("notification-service: order $orderId's confirmation email dead-lettered: ${exception.message}")
}
}

A dead-lettered message is never silently dropped — it’s durably recorded in a queryable table and surfaces as an alert, giving Northwind’s support team a concrete, actionable record (and, notably, a natural trigger to reach out to the customer through another channel if their email address turns out to be wrong).

6. Step-by-Step Flow

DeadLetterAlertServicedead-letter topicretry topic (5m)retry topic (30s)OrderConfirmationListenerorder-confirmationsDeadLetterAlertServicedead-letter topicretry topic (5m)retry topic (30s)OrderConfirmationListenerorder-confirmationswait 30 secondswait 5 minutesorder A done — no manual intervention neededOrderPlacedNotification (order A)send fails: TransientProviderFailureroute to 30s retry topicreattemptsend fails again: TransientProviderFailureroute to 5m retry topicreattemptsend succeedsOrderPlacedNotification (order B, invalid address)send fails: PermanentDeliveryFailureroute directly to DLQ (excluded from retry)recordAndAlert(order B)
DeadLetterAlertServicedead-letter topicretry topic (5m)retry topic (30s)OrderConfirmationListenerorder-confirmationsDeadLetterAlertServicedead-letter topicretry topic (5m)retry topic (30s)OrderConfirmationListenerorder-confirmationswait 30 secondswait 5 minutesorder A done — no manual intervention neededOrderPlacedNotification (order A)send fails: TransientProviderFailureroute to 30s retry topicreattemptsend fails again: TransientProviderFailureroute to 5m retry topicreattemptsend succeedsOrderPlacedNotification (order B, invalid address)send fails: PermanentDeliveryFailureroute directly to DLQ (excluded from retry)recordAndAlert(order B)
  1. Client action. Two orders, A and B, are placed around the same time; order A’s customer has a valid but temporarily unreachable email provider, order B’s customer mistyped their address.
  2. API request equivalent. Both OrderPlacedNotification events arrive at notification-service’s listener.
  3. Service behavior. Order A’s send throws TransientProviderFailure; order B’s throws PermanentDeliveryFailure — the classification made at the point of failure, based on the actual response from the email provider’s API.
  4. Database interaction. sentEmailLog is checked and updated identically to Chapter 18, once a send finally succeeds.
  5. Inter-service communication. No other service is involved — this pattern operates entirely within notification-service’s own message-processing pipeline.
  6. Error or failure handling. Order A’s message is retried automatically after 30 seconds, then (if still failing) after 5 minutes, eventually succeeding once the provider recovers — never blocking the main topic’s other messages while waiting. Order B’s message skips retries entirely and lands in the DLQ immediately, since no amount of waiting will fix a malformed address.
  7. Observability signals. Track messages-per-hour landing in the DLQ (a rising rate signals either a systemic email-quality problem worth investigating at checkout, or a broader provider issue) separately from retry-topic depth (a rising retry backlog signals an ongoing provider outage still in progress).
  8. Final response/outcome. Order A’s confirmation eventually arrives, automatically, with no human involvement. Order B’s failure is durably recorded and alerted, giving support a concrete, actionable item — a materially better outcome than either blocking every other pending confirmation behind order B, or silently losing its notification.

7. Production Concerns

  • Timeouts, retries, idempotency. Every retry stage reprocesses the message from scratch — sentEmailLog.alreadySent() must be checked on every attempt, including the first retry after a partial failure, or a message that succeeded but failed to record its success could send a duplicate.
  • Data consistency. The dead-letter record and the alert should be written durably before considering the message “handled” — an alert sent but never persisted to dead_lettered_notifications leaves no queryable record once the alert scrolls out of whatever channel it was sent to.
  • API versioning and backward compatibility. Retry and dead-letter topics carry the same message schema as the main topic — no new versioning surface, but note that a schema change must stay compatible with messages that might still be sitting in a retry topic’s delay window from before the change shipped.
  • Authentication and service-to-service trust. Retry and DLQ topics need the same ACL scoping as the main topic (Chapter 7) — don’t leave them more permissive by oversight just because they’re “internal plumbing.”
  • Logging, metrics, tracing, correlation IDs. Propagate the original correlation ID through every retry attempt and into the dead-letter record — a support engineer investigating a dead-lettered message needs to trace it back to the original order and request without guesswork.
  • Kubernetes deployment. No new infrastructure beyond additional Kafka topics — the retry and DLQ consumers can run inside the same notification-service deployment (as this chapter’s @RetryableTopic annotation does automatically) rather than as separate services, unless retry/DLQ volume grows large enough to justify isolating them.
  • Testing strategy. Test the classification logic (which exceptions route to retry vs. straight to DLQ) as its own unit test, independent of the Kafka machinery — then a smaller number of integration tests confirming the actual topic routing behaves as configured.
  • Migration strategy. Introduce dead-lettering first, even without a retry chain, for any consumer that currently has no failure-handling story beyond default framework behavior — a DLQ alone is a meaningful improvement over silent blocking or silent loss; add retry stages once a real transient-failure pattern justifies the added delay chain.

8. Common Mistakes

  1. Not distinguishing transient from permanent failures. Retrying a malformed email address three times, at increasing delays, wastes eight-plus minutes before finally giving up on something that could have been recognized as unrecoverable instantly. Fix: classify failures explicitly, as Section 5’s EmailSendException hierarchy does, and skip retries entirely for known-permanent cases.
  2. A DLQ nobody monitors. Routing failed messages to a dead-letter topic satisfies the “don’t block the main partition” requirement but not the “don’t silently lose work” one, if nothing ever alerts on or reviews what lands there. Fix: pair every DLQ with an active alert and a reviewable record, as Section 5’s DeadLetterAlertService does — never treat the DLQ as a place messages go to be forgotten.
  3. Retrying the main topic’s own partition instead of using a separate retry topic. Configuring a high retry count directly on the main consumer, rather than routing to separate retry topics, still blocks that partition’s other messages for the full retry duration, since the offset can’t advance past an unresolved message. Fix: use a genuinely separate retry topic (or topic chain), as Spring Kafka’s @RetryableTopic does, so the main topic’s consumer keeps moving.
  4. Retry delays not calibrated to the actual failure’s typical recovery time. A 1-second retry delay for a provider outage that typically takes minutes to resolve just burns through the retry budget uselessly fast; an unnecessarily long delay for a failure that usually self-resolves in seconds delays legitimate deliveries. Fix: calibrate backoff stages to the observed recovery pattern of the actual failure mode, as Northwind’s 30s-then-5min chain is meant to reflect its provider’s typical outage duration.
  5. Non-idempotent processing combined with retries. If sendConfirmation had a side effect that wasn’t safe to repeat (charging a fee for sending, say, in some hypothetical provider), retrying it could compound that side effect. Fix: verify idempotency before enabling retries for any workload, exactly as Chapter 1’s ongoing discipline requires.
  6. Treating a growing DLQ as acceptable background noise. Once a DLQ exists, it’s tempting to let a slow trickle of dead-lettered messages accumulate without addressing the root cause (a checkout form with no email validation, say) that keeps producing them. Fix: treat a sustained non-zero DLQ rate as a signal to fix the upstream cause, not just a permanent, tolerated cost of doing business.

9. Decision Guide

Problem signalUse this pattern?WhyAlternative
Non-trivial failure rate with a real transient/permanent distinctionYes (retry topics + DLQ)Retries transient failures automatically; surfaces permanent ones immediately for review—
A poison message currently blocks an entire partition’s other workYesRouting failures to a separate topic unblocks the main consumer—
Every failure in the workload is permanent, none are transientDLQ only, no retry chainNo failure mode benefits from a delay-and-retry; skip straight to dead-letteringSimple DLQ with immediate routing
Very low, rare failure rate, currently handled fine manuallyNot yetDedicated infrastructure isn’t justified below a certain failure frequencyManual handling with basic alerting
Processing isn’t idempotent and can’t easily be made soFix idempotency firstRetrying non-idempotent work risks compounding a side effectAddress idempotency before adopting retry topics

10. Hands-On Exercise

Extend it: add a periodic job that queries dead_lettered_notifications for entries older than 24 hours with no resolution recorded, and escalates them to a higher-priority alert — ensuring dead-lettered messages don’t just sit reviewed-but-unresolved indefinitely.

Simulate a failure: configure a test email sender that always throws TransientProviderFailure for the first two attempts and succeeds on the third, and confirm the message is retried through both backoff stages and ultimately succeeds — then repeat with a sender that always throws PermanentDeliveryFailure and confirm it skips straight to the DLQ.

Decision question, with justification required: Northwind’s DLQ has been steadily accumulating about fifty dead-lettered emails per day, almost all with the failure reason “recipient mailbox does not exist.” Is more retry-topic tuning the right response, or does this number point to a completely different fix, upstream of notification-service entirely? Name where in the system you’d actually make a change, and why.

11. Key Takeaways

  • A dead-letter topic preserves failed messages for review instead of silently dropping them or blocking a partition’s other work indefinitely — pair it with active alerting, or it’s equivalent to silent data loss with extra steps.
  • A retry topic (or a short chain of them) lets transient failures be reattempted after a calibrated delay without blocking the main topic’s consumer while waiting.
  • The single most important design decision is classifying failures as transient or permanent — retrying a permanent failure wastes the retry budget; failing to retry a transient one gives up too early.
  • Calibrate retry delays to the actual, observed recovery time of the failure modes you’re handling, not to an arbitrary default.
  • This pattern requires the same idempotency discipline this series has required since Chapter 1 — a message may be reprocessed after a retry delay, and that reprocessing must be safe.
  • A sustained, non-trivial DLQ rate is a signal to fix the upstream cause (bad input validation, a flaky integration) — not a permanent, tolerated cost to be monitored and ignored indefinitely.
  • This pattern and Chapter 18’s competing consumers work together naturally: competing consumers scale throughput for healthy messages; dead-letter and retry topics ensure a failing message never becomes the bottleneck that throughput scaling was meant to prevent.
Spring BootKotlinMicroservicesKafka

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind