The dead-letter queue alarm fires, or someone notices it holds 40,000 messages instead of the usual dozen. The instinct is to hit “redrive” and move on. That usually sends the same messages back through the same failure and fills the DLQ again within the hour. This guide is the order of operations: what to look at first, how to find out why messages failed, and when a redrive is actually safe.
A dead-letter queue that keeps growing means messages are failing faster than anyone handles them. Do not redrive first. Find the failure class, fix it, then replay under control:
- Read the shape: a step at a deploy, a steady ramp, a burst during an outage, or growth that tracks load.
- Sample without deleting, and join each message to its error in logs, traces, or DLQ headers.
- Classify: poison message, code or schema bug, dependency outage, retry or lease misconfiguration, or missing prerequisite.
- Redrive only after the fix, with an idempotent consumer, a small test batch, and a rate limit.
The steps below apply to Amazon SQS, RabbitMQ, and Kafka dead-letter topics. The mechanics differ, and they matter, because each system decides differently what counts as “failed” and records different amounts of information about why.
What does a growing dead-letter queue mean?
A dead-letter queue (DLQ) is a separate queue or topic that holds messages the consumer could not process after a set number of attempts, so they stop blocking or looping on the main queue. A DLQ that grows means the rate of failed messages exceeds the rate at which anyone resolves them. That is always a symptom. The DLQ tells you that messages failed; it rarely tells you why.
Before investigating, be precise about what your system counts as a failure:
| System | What sends a message to the DLQ | What is recorded about why |
|---|---|---|
| Amazon SQS | Redrive policy: the receive count exceeds maxReceiveCount | Nothing. The message is moved as-is; the reason lives only in your consumer’s logs. |
| RabbitMQ | Dead-letter exchange: rejected or nacked with requeue=false, TTL expired, queue length limit exceeded, or a quorum queue delivery limit exceeded | An x-death header with the reason, source queue, and count |
| Spring for Apache Kafka | DeadLetterPublishingRecoverer after the error handler’s retries are exhausted | Headers with the exception class, message, stack trace, and original topic, partition, and offset |
| Kafka Connect (sink) | errors.tolerance=all and a dead-letter topic configured | Error context headers, when context headers are enabled |
Kafka Connect settings: errors.deadletterqueue.topic.name and errors.deadletterqueue.context.headers.enable=true.
SQS needs a closer look, because its rule surprises people. SQS does not count failures. It counts receives. Every time a consumer receives a message and does not delete it before the visibility timeout expires, the message becomes visible again with a higher receive count. When that count passes maxReceiveCount, SQS moves it to the DLQ. The SQS dead-letter queue documentation covers the redrive policy in full.
{
"deadLetterTargetArn": "arn:aws:sqs:us-east-1:123456789012:orders-dlq",
"maxReceiveCount": "5"
}
Is it a step, a ramp, or a burst?
Start with the shape of the growth, because it narrows the cause before you read a single message. Plot DLQ depth and the age of the oldest message over the last few days, then line them up with deploys, config changes, and dependency incidents. For SQS, the CloudWatch metrics are ApproximateNumberOfMessagesVisible and ApproximateAgeOfOldestMessage on the DLQ. Watch depth rather than NumberOfMessagesSent: AWS documents that messages moved to a DLQ by the redrive policy are not counted in the DLQ’s NumberOfMessagesSent.
| Signal | What it usually means |
|---|---|
| Growth starts within minutes of a deploy | The release broke parsing, validation, or a downstream call. Find the commit before touching the backlog. |
| Growth starts at a producer’s deploy, not yours | A payload or schema change upstream. Compare message versions before and after. |
| Steady trickle for weeks | A small class of messages always fails: an edge case, an old client, a tenant with unusual data. Nobody noticed because the DLQ had no alarm. |
| Burst that matches a dependency incident | Retries were exhausted during the outage. The messages are probably fine now. |
| Growth only at peak traffic | Processing slows under load and outlives the visibility timeout, or retries are too few for normal throttling. |
| Oldest message age keeps rising while depth is flat | Nobody is draining the DLQ at all. The oldest messages will expire at the end of the retention period. |
How do you inspect DLQ messages without making it worse?
Inspect a sample of messages without deleting anything, and join each one to the error that sent it there. Two mistakes are common here: consuming the DLQ with a script that deletes as it reads, and assuming the DLQ itself knows why the message failed.
In SQS, reading a message is a receive. It hides the message for the visibility timeout and increases its receive count, and the console’s “poll for messages” is also a receive. That is harmless on a DLQ with no redrive policy of its own, but do not point an inspection script at the source queue during an incident, and never delete during sampling. A small read-only sampler is enough:
import boto3, collections, json
sqs = boto3.client("sqs")
seen, groups = {}, collections.Counter()
for _ in range(20):
resp = sqs.receive_message(
QueueUrl=DLQ_URL, MaxNumberOfMessages=10, WaitTimeSeconds=1,
VisibilityTimeout=600, # hide sampled messages so we don't re-read them
MessageSystemAttributeNames=["All"], MessageAttributeNames=["All"],
)
for m in resp.get("Messages", []):
body = json.loads(m["Body"])
seen[m["MessageId"]] = body.get("event_id") # ids to look up in logs
groups[(body.get("type"), body.get("schema_version"))] += 1
# no delete_message: the DLQ stays intact
for key, n in groups.most_common():
print(n, key)
For a Kafka dead-letter topic, read with a new consumer group or kafka-console-consumer.sh --from-beginning --property print.headers=true. Reading does not remove records, and the Spring headers such as kafka_dlt-exception-fqcn, kafka_dlt-exception-message, and kafka_dlt-original-offset give you the reason directly. For RabbitMQ, the x-death header on each dead-lettered message shows the reason: rejected, expired, maxlen, or delivery_limit.
Join each message to its failure
With SQS, the reason lives in the consumer’s logs and traces, not in the queue. Search for the sampled ids around each message’s receive attempts and collect the exception and stack trace. This only works if the consumer logs a message id or business id with every failure, which is the first thing to fix if it does not. How to Correlate Logs, Traces, and a Code Change covers doing this join quickly under pressure.
Then group by error signature: the exception type plus the top application frame, with ids and values stripped out. In most incidents, one or two signatures account for almost all of the backlog. Fix those, and treat the long tail separately.
A poison message is a message that fails every time it is processed, no matter how often it is retried, because of its content: invalid data, an unsupported version, or a reference to something that will never exist. Retrying a poison message only wastes consumer capacity until it is dead-lettered.
Which kind of failure is it?
Every dead-lettered message falls into one of five classes, and each class needs a different action before any redrive. Classify by evidence, not by the loudest error in the logs.
| Class | Evidence | Fix before redrive | Redrive? |
|---|---|---|---|
| Poison message | Same message fails on every attempt; validation or parse errors; one tenant or client | Handle the case in code, or reject it as non-retryable with a recorded reason | Only after the code handles it; otherwise archive |
| Code or schema bug | Step at a deploy; one error signature dominates; producer version changed | Fix or roll back the consumer, or make it accept both versions | Yes, after the fix is live |
| Dependency outage | Timeouts, 5xx, or throttling from one dependency, bounded to its incident window | None in the consumer; review retry and backoff settings | Yes, rate-limited |
| Lease or retry settings | Logs show success or long processing; receive counts climb without errors | Raise the visibility timeout, add a heartbeat, adjust maxReceiveCount | Carefully; some already succeeded |
| Missing prerequisite | “Not found” errors for entities created moments later; ordering assumptions | Delayed retry, ordering by key, or tolerate and reconcile | Yes, once the entity exists |
The lease class deserves attention because it hides. When processing gets slower, for example after a dependency slows down or a batch size increases, messages that were going to succeed are received again and again until SQS dead-letters them. Some of them have already been processed, possibly more than once. Redriving those messages processes them yet again, so it is only safe with an idempotent consumer. SQS or Kafka Message Processed Twice: Designing Safe Consumers covers how to build one.
A dead-letter queue is not a place where failures are handled. It is where they wait for someone to handle them.
When is it safe to redrive?
Redrive is safe when four things are true: the cause is fixed and deployed, the consumer is idempotent, the messages are still valid to apply, and the replay rate will not overload the consumer or its dependencies. Check each one explicitly.
Gate 3 is the one teams skip. A message that sat in the DLQ for three days may describe a state that is no longer true: a price that has since changed, an address that was updated, a cancellation that arrived after the original order. Replaying it can overwrite newer data. Consumers that apply events with a version or timestamp check (“only update if this event is newer than what we have”) are safe here; consumers that blindly set values are not.
For SQS, use the console’s DLQ redrive or the StartMessageMoveTask API. Without a destination ARN, messages go back to the queue they came from. Set a rate:
# Current backlog (queue attributes, not CloudWatch metric names)
aws sqs get-queue-attributes --queue-url "$DLQ_URL" \
--attribute-names ApproximateNumberOfMessages ApproximateNumberOfMessagesNotVisible
# Move messages back to their source queue at 10 per second
aws sqs start-message-move-task \
--source-arn arn:aws:sqs:us-east-1:123456789012:orders-dlq \
--max-number-of-messages-per-second 10
# Watch progress, or stop it if errors return
aws sqs list-message-move-tasks --source-arn arn:aws:sqs:us-east-1:123456789012:orders-dlq
aws sqs cancel-message-move-task --task-handle "$TASK_HANDLE"
For a canary, move a small number of messages first. With SQS you can run a short move task and cancel it, or receive a handful from the DLQ with a script and send them to the source queue. For Kafka, there is no built-in redrive: a small replay consumer reads the dead-letter topic and republishes to the original topic, or processes the records directly, with its own offsets so you can stop and resume.
For SQS standard queues, a message keeps its original enqueue timestamp when it moves to the DLQ, so it expires based on when it was first sent, not when it was dead-lettered. AWS recommends setting the DLQ’s retention period longer than the source queue’s. If both use the 4-day default, a message that took three days to reach the DLQ has one day left. Check ApproximateAgeOfOldestMessage against the retention period before planning a slow investigation.
How do you stop the DLQ from growing again?
Most DLQs that grow unnoticed have three gaps: no alarm, no owner, and no distinction between errors worth retrying and errors that never will succeed. Close all three.
- Alarm on depth and age. For most queues, one message in the DLQ is worth a ticket, and a rising oldest-message age is worth a page before retention runs out.
- Separate retryable from non-retryable errors. Validation and parse failures will not improve with retries. Send them straight to the DLQ with the reason attached, and retry only transient errors with backoff. Blanket retries on everything are how a small bug becomes a retry storm.
- Carry the reason with the message. Spring and RabbitMQ do this with headers. For SQS, a consumer that dead-letters a message itself can add the error as a message attribute.
- Size the lease and the retry count together. A visibility timeout above worst-case processing, and a
maxReceiveCountthat tolerates a short blip. - Give the DLQ an owner and a runbook, including a redrive procedure that is safe to run twice. The Production Readiness Checklist for a Backend Service lists this with the other runbook items.
@Bean
DefaultErrorHandler errorHandler(KafkaTemplate<Object, Object> template) {
// Publishes failed records to <topic>.DLT (same partition) with exception headers
var recoverer = new DeadLetterPublishingRecoverer(template);
var handler = new DefaultErrorHandler(recoverer, new FixedBackOff(1000L, 3L));
handler.addNotRetryableExceptions(ValidationException.class, JsonParseException.class);
return handler;
}
Because the default recoverer writes to the same partition number on the .DLT topic, create the dead-letter topic with at least as many partitions as the source. The same pattern, separating the error classes and recording why, applies to any framework.
How Tomosu helps
Tomosu analyzes a repository and its pull requests for production reliability risks. Most of what makes a DLQ grow is visible in code before it is visible in a dashboard: how a consumer classifies errors, what it retries, where it acknowledges, and what it logs when it fails. For this problem, Tomosu:
- Maps consumers and their failure paths: which exceptions are retried, which are dead-lettered, and which are swallowed.
- Flags handlers that will not be safe to redrive, such as non-idempotent side effects or state updates with no version check.
- Points at changes that shift the failure rate: new validation, a changed payload contract, a slower downstream call inside the consumer, or a reduced timeout.
- Surfaces missing failure context, such as error logging without a message or business id, which makes the join in this guide impossible.
These signals feed the Production Reliability Index. To try it on your consumers, follow the repository scan guide.
Scan your repository with Tomosu →
Key takeaways
- A dead-letter queue that keeps growing is a symptom; do not redrive before you know the cause.
- Read the growth shape first: a step at a deploy, a steady ramp, a burst during an outage, or growth that tracks traffic.
- SQS counts receives, not failures, and stores no reason; join messages to consumer logs by id.
- Sample without deleting, then group failures by error signature.
- Classify each group as poison, bug, outage, lease settings, or missing prerequisite, and fix accordingly.
- Redrive only through four gates: cause fixed, consumer idempotent, messages still valid, canary batch clean; then rate-limit.
- Prevent the next backlog with depth and age alarms, non-retryable error routing, failure context, and an owner.
Frequently asked questions
Why does my dead-letter queue keep growing?
Messages are failing faster than anyone handles them. The usual causes are a code or schema bug released in a deploy, a dependency outage that exhausted retries, a slice of malformed or unexpected messages, or retry settings such as a visibility timeout or maxReceiveCount that dead-letter messages which would have succeeded. The growth pattern and a sample of messages tell you which.
Should I redrive the DLQ right away?
No. Redriving before the cause is fixed sends the same messages back through the same failure, and they return to the DLQ, usually with extra load on the consumer and its dependencies. Fix the cause first, confirm the consumer is idempotent and the messages are still valid, then redrive a small batch before the rest.
Why do messages that succeeded end up in the SQS dead-letter queue?
SQS moves a message after its receive count exceeds maxReceiveCount, and it counts receives, not failures. If processing takes longer than the visibility timeout, or the consumer never deletes the message, it is received again even though the work succeeded. Viewing messages in the console also receives them and increases the count.
How do I see why a message was dead-lettered in SQS?
SQS does not store the failure reason with the message. You have to join the message to the consumer’s own record of the failure, usually logs or traces keyed by the message id or a business id in the payload. Consumers that dead-letter messages themselves can add the reason as a message attribute.
How do I redrive an SQS dead-letter queue?
Use the console’s DLQ redrive or the StartMessageMoveTask API, for example aws sqs start-message-move-task with the DLQ as the source ARN. Without a destination ARN, messages go back to their original source queue. Set a maximum number of messages per second so the redrive does not flood the consumer.
Does Kafka have a dead-letter queue?
Not in the broker. Dead-letter topics are implemented by clients and frameworks. Spring for Apache Kafka’s DeadLetterPublishingRecoverer publishes failed records to a topic named after the original with a .DLT suffix by default, and Kafka Connect sink connectors can route failures to a topic set in errors.deadletterqueue.topic.name.
What maxReceiveCount should I use?
High enough that a short dependency blip or a single slow attempt does not dead-letter healthy messages, and low enough that a poison message does not occupy consumers for long. Values in the single digits to low tens are common. Pair it with a visibility timeout above your worst-case processing time.
Every message in a dead-letter queue is a failure that already happened. Tomosu helps you see the consumer paths that will put the next ones there, before they ship. Assess your repository →