Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Queues

Dead-Letter Queue Keeps Growing: Where to Start

Tomosu AI·13 min read·

The dead-letter queue alarm fires, or someone notices it holds 40,000 messages instead of the usual dozen. The instinct is to hit “redrive” and move on. That usually sends the same messages back through the same failure and fills the DLQ again within the hour. This guide is the order of operations: what to look at first, how to find out why messages failed, and when a redrive is actually safe.

Quick answer

A dead-letter queue that keeps growing means messages are failing faster than anyone handles them. Do not redrive first. Find the failure class, fix it, then replay under control:

The steps below apply to Amazon SQS, RabbitMQ, and Kafka dead-letter topics. The mechanics differ, and they matter, because each system decides differently what counts as “failed” and records different amounts of information about why.

What does a growing dead-letter queue mean?

A dead-letter queue (DLQ) is a separate queue or topic that holds messages the consumer could not process after a set number of attempts, so they stop blocking or looping on the main queue. A DLQ that grows means the rate of failed messages exceeds the rate at which anyone resolves them. That is always a symptom. The DLQ tells you that messages failed; it rarely tells you why.

Before investigating, be precise about what your system counts as a failure:

SystemWhat sends a message to the DLQWhat is recorded about why
Amazon SQSRedrive policy: the receive count exceeds maxReceiveCountNothing. The message is moved as-is; the reason lives only in your consumer’s logs.
RabbitMQDead-letter exchange: rejected or nacked with requeue=false, TTL expired, queue length limit exceeded, or a quorum queue delivery limit exceededAn x-death header with the reason, source queue, and count
Spring for Apache KafkaDeadLetterPublishingRecoverer after the error handler’s retries are exhaustedHeaders with the exception class, message, stack trace, and original topic, partition, and offset
Kafka Connect (sink)errors.tolerance=all and a dead-letter topic configuredError context headers, when context headers are enabled

Kafka Connect settings: errors.deadletterqueue.topic.name and errors.deadletterqueue.context.headers.enable=true.

SQS needs a closer look, because its rule surprises people. SQS does not count failures. It counts receives. Every time a consumer receives a message and does not delete it before the visibility timeout expires, the message becomes visible again with a higher receive count. When that count passes maxReceiveCount, SQS moves it to the DLQ. The SQS dead-letter queue documentation covers the redrive policy in full.

HOW A MESSAGE REACHES AN SQS DLQ Source queue orders receive Consumer receive count += 1 DeleteMessage Done removed from queue Not deleted in time throws, times out, crashes, or succeeds too slowly visible again after the visibility timeout count > max (5) Dead-letter queue orders-dlq SQS counts receives, not failures. A slow success, a missing delete, or someone polling in the console all add to the count. No error reason travels with the message.
In SQS, the DLQ is a count threshold. Messages can arrive there without ever throwing an error.
SQS · RedrivePolicy attribute on the source queue
{
  "deadLetterTargetArn": "arn:aws:sqs:us-east-1:123456789012:orders-dlq",
  "maxReceiveCount": "5"
}

Is it a step, a ramp, or a burst?

Start with the shape of the growth, because it narrows the cause before you read a single message. Plot DLQ depth and the age of the oldest message over the last few days, then line them up with deploys, config changes, and dependency incidents. For SQS, the CloudWatch metrics are ApproximateNumberOfMessagesVisible and ApproximateAgeOfOldestMessage on the DLQ. Watch depth rather than NumberOfMessagesSent: AWS documents that messages moved to a DLQ by the redrive policy are not counted in the DLQ’s NumberOfMessagesSent.

DLQ DEPTH OVER TIME: FOUR SHAPES, FOUR SUSPECTS Step at a deploy deploy v142 Suspect: code, schema, or config change Slow, steady ramp Suspect: a slice of poison or unexpected input Burst, then flat dependency down Suspect: an outage that exhausted retries Tracks traffic traffic (dashed) Suspect: timeouts or leases too short at peak
The shape of the growth points at a suspect before you read any messages. Always overlay deploys and traffic.
SignalWhat it usually means
Growth starts within minutes of a deployThe release broke parsing, validation, or a downstream call. Find the commit before touching the backlog.
Growth starts at a producer’s deploy, not yoursA payload or schema change upstream. Compare message versions before and after.
Steady trickle for weeksA small class of messages always fails: an edge case, an old client, a tenant with unusual data. Nobody noticed because the DLQ had no alarm.
Burst that matches a dependency incidentRetries were exhausted during the outage. The messages are probably fine now.
Growth only at peak trafficProcessing slows under load and outlives the visibility timeout, or retries are too few for normal throttling.
Oldest message age keeps rising while depth is flatNobody is draining the DLQ at all. The oldest messages will expire at the end of the retention period.

How do you inspect DLQ messages without making it worse?

Inspect a sample of messages without deleting anything, and join each one to the error that sent it there. Two mistakes are common here: consuming the DLQ with a script that deletes as it reads, and assuming the DLQ itself knows why the message failed.

In SQS, reading a message is a receive. It hides the message for the visibility timeout and increases its receive count, and the console’s “poll for messages” is also a receive. That is harmless on a DLQ with no redrive policy of its own, but do not point an inspection script at the source queue during an incident, and never delete during sampling. A small read-only sampler is enough:

Python · sample the DLQ, delete nothingread-only
import boto3, collections, json

sqs = boto3.client("sqs")
seen, groups = {}, collections.Counter()

for _ in range(20):
    resp = sqs.receive_message(
        QueueUrl=DLQ_URL, MaxNumberOfMessages=10, WaitTimeSeconds=1,
        VisibilityTimeout=600,                 # hide sampled messages so we don't re-read them
        MessageSystemAttributeNames=["All"], MessageAttributeNames=["All"],
    )
    for m in resp.get("Messages", []):
        body = json.loads(m["Body"])
        seen[m["MessageId"]] = body.get("event_id")     # ids to look up in logs
        groups[(body.get("type"), body.get("schema_version"))] += 1
        # no delete_message: the DLQ stays intact

for key, n in groups.most_common():
    print(n, key)

For a Kafka dead-letter topic, read with a new consumer group or kafka-console-consumer.sh --from-beginning --property print.headers=true. Reading does not remove records, and the Spring headers such as kafka_dlt-exception-fqcn, kafka_dlt-exception-message, and kafka_dlt-original-offset give you the reason directly. For RabbitMQ, the x-death header on each dead-lettered message shows the reason: rejected, expired, maxlen, or delivery_limit.

Join each message to its failure

With SQS, the reason lives in the consumer’s logs and traces, not in the queue. Search for the sampled ids around each message’s receive attempts and collect the exception and stack trace. This only works if the consumer logs a message id or business id with every failure, which is the first thing to fix if it does not. How to Correlate Logs, Traces, and a Code Change covers doing this join quickly under pressure.

Then group by error signature: the exception type plus the top application frame, with ids and values stripped out. In most incidents, one or two signatures account for almost all of the backlog. Fix those, and treat the long tail separately.

Definition

A poison message is a message that fails every time it is processed, no matter how often it is retried, because of its content: invalid data, an unsupported version, or a reference to something that will never exist. Retrying a poison message only wastes consumer capacity until it is dead-lettered.

Which kind of failure is it?

Every dead-lettered message falls into one of five classes, and each class needs a different action before any redrive. Classify by evidence, not by the loudest error in the logs.

WHY WAS THIS MESSAGE DEAD-LETTERED? Q1 · REPLAY ONE BY HAND Does the same message fail every time? Q2 · DEPLOY TIMELINE Did failures start with a deploy (yours or a producer’s)? Q3 · DEPENDENCY HEALTH Do errors match an outage or throttling window? Q4 · CONSUMER LOGS Did it succeed, or time out, before dead-lettering? Poison message Fix validation, or drop with a record Code or schema bug Fix and deploy, then redrive Dependency outage Redrive after it recovers Lease or retry settings Visibility timeout, maxReceiveCount YESYESYESYES NONONONO Otherwise, missing prerequisite: it arrived before the entity it refers to. Add ordering or delayed retry.
Classify before you act. Only one of the five classes is usually safe to redrive as it is.
ClassEvidenceFix before redriveRedrive?
Poison messageSame message fails on every attempt; validation or parse errors; one tenant or clientHandle the case in code, or reject it as non-retryable with a recorded reasonOnly after the code handles it; otherwise archive
Code or schema bugStep at a deploy; one error signature dominates; producer version changedFix or roll back the consumer, or make it accept both versionsYes, after the fix is live
Dependency outageTimeouts, 5xx, or throttling from one dependency, bounded to its incident windowNone in the consumer; review retry and backoff settingsYes, rate-limited
Lease or retry settingsLogs show success or long processing; receive counts climb without errorsRaise the visibility timeout, add a heartbeat, adjust maxReceiveCountCarefully; some already succeeded
Missing prerequisite“Not found” errors for entities created moments later; ordering assumptionsDelayed retry, ordering by key, or tolerate and reconcileYes, once the entity exists

The lease class deserves attention because it hides. When processing gets slower, for example after a dependency slows down or a batch size increases, messages that were going to succeed are received again and again until SQS dead-letters them. Some of them have already been processed, possibly more than once. Redriving those messages processes them yet again, so it is only safe with an idempotent consumer. SQS or Kafka Message Processed Twice: Designing Safe Consumers covers how to build one.

A dead-letter queue is not a place where failures are handled. It is where they wait for someone to handle them.

When is it safe to redrive?

Redrive is safe when four things are true: the cause is fixed and deployed, the consumer is idempotent, the messages are still valid to apply, and the replay rate will not overload the consumer or its dependencies. Check each one explicitly.

FOUR GATES BEFORE A REDRIVE GATE 1 Cause fixed and deployed GATE 2 Consumer is idempotent GATE 3 Messages still valid to apply GATE 4 Canary batch succeeds Redrive the rest, rate-limited no: messagesfail again no: replaysduplicate effects no: stale eventsoverwrite state no: stop, lookat new errors During the redrive, watch DLQ depth, consumer error rate, and load on the dependencies the consumer calls.
A redrive is a deploy of old traffic. Gate it the way you would gate a release.

Gate 3 is the one teams skip. A message that sat in the DLQ for three days may describe a state that is no longer true: a price that has since changed, an address that was updated, a cancellation that arrived after the original order. Replaying it can overwrite newer data. Consumers that apply events with a version or timestamp check (“only update if this event is newer than what we have”) are safe here; consumers that blindly set values are not.

For SQS, use the console’s DLQ redrive or the StartMessageMoveTask API. Without a destination ARN, messages go back to the queue they came from. Set a rate:

AWS CLI · rate-limited redrive
# Current backlog (queue attributes, not CloudWatch metric names)
aws sqs get-queue-attributes --queue-url "$DLQ_URL" \
  --attribute-names ApproximateNumberOfMessages ApproximateNumberOfMessagesNotVisible

# Move messages back to their source queue at 10 per second
aws sqs start-message-move-task \
  --source-arn arn:aws:sqs:us-east-1:123456789012:orders-dlq \
  --max-number-of-messages-per-second 10

# Watch progress, or stop it if errors return
aws sqs list-message-move-tasks --source-arn arn:aws:sqs:us-east-1:123456789012:orders-dlq
aws sqs cancel-message-move-task --task-handle "$TASK_HANDLE"

For a canary, move a small number of messages first. With SQS you can run a short move task and cancel it, or receive a handful from the DLQ with a script and send them to the source queue. For Kafka, there is no built-in redrive: a small replay consumer reads the dead-letter topic and republishes to the original topic, or processes the records directly, with its own offsets so you can stop and resume.

Check the DLQ retention period

For SQS standard queues, a message keeps its original enqueue timestamp when it moves to the DLQ, so it expires based on when it was first sent, not when it was dead-lettered. AWS recommends setting the DLQ’s retention period longer than the source queue’s. If both use the 4-day default, a message that took three days to reach the DLQ has one day left. Check ApproximateAgeOfOldestMessage against the retention period before planning a slow investigation.

How do you stop the DLQ from growing again?

Most DLQs that grow unnoticed have three gaps: no alarm, no owner, and no distinction between errors worth retrying and errors that never will succeed. Close all three.

Java · Spring for Apache Kafkaretry transient, dead-letter poison
@Bean
DefaultErrorHandler errorHandler(KafkaTemplate<Object, Object> template) {
    // Publishes failed records to <topic>.DLT (same partition) with exception headers
    var recoverer = new DeadLetterPublishingRecoverer(template);
    var handler = new DefaultErrorHandler(recoverer, new FixedBackOff(1000L, 3L));
    handler.addNotRetryableExceptions(ValidationException.class, JsonParseException.class);
    return handler;
}

Because the default recoverer writes to the same partition number on the .DLT topic, create the dead-letter topic with at least as many partitions as the source. The same pattern, separating the error classes and recording why, applies to any framework.

How Tomosu helps

Tomosu analyzes a repository and its pull requests for production reliability risks. Most of what makes a DLQ grow is visible in code before it is visible in a dashboard: how a consumer classifies errors, what it retries, where it acknowledges, and what it logs when it fails. For this problem, Tomosu:

These signals feed the Production Reliability Index. To try it on your consumers, follow the repository scan guide.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

Why does my dead-letter queue keep growing?

Messages are failing faster than anyone handles them. The usual causes are a code or schema bug released in a deploy, a dependency outage that exhausted retries, a slice of malformed or unexpected messages, or retry settings such as a visibility timeout or maxReceiveCount that dead-letter messages which would have succeeded. The growth pattern and a sample of messages tell you which.

Should I redrive the DLQ right away?

No. Redriving before the cause is fixed sends the same messages back through the same failure, and they return to the DLQ, usually with extra load on the consumer and its dependencies. Fix the cause first, confirm the consumer is idempotent and the messages are still valid, then redrive a small batch before the rest.

Why do messages that succeeded end up in the SQS dead-letter queue?

SQS moves a message after its receive count exceeds maxReceiveCount, and it counts receives, not failures. If processing takes longer than the visibility timeout, or the consumer never deletes the message, it is received again even though the work succeeded. Viewing messages in the console also receives them and increases the count.

How do I see why a message was dead-lettered in SQS?

SQS does not store the failure reason with the message. You have to join the message to the consumer’s own record of the failure, usually logs or traces keyed by the message id or a business id in the payload. Consumers that dead-letter messages themselves can add the reason as a message attribute.

How do I redrive an SQS dead-letter queue?

Use the console’s DLQ redrive or the StartMessageMoveTask API, for example aws sqs start-message-move-task with the DLQ as the source ARN. Without a destination ARN, messages go back to their original source queue. Set a maximum number of messages per second so the redrive does not flood the consumer.

Does Kafka have a dead-letter queue?

Not in the broker. Dead-letter topics are implemented by clients and frameworks. Spring for Apache Kafka’s DeadLetterPublishingRecoverer publishes failed records to a topic named after the original with a .DLT suffix by default, and Kafka Connect sink connectors can route failures to a topic set in errors.deadletterqueue.topic.name.

What maxReceiveCount should I use?

High enough that a short dependency blip or a single slow attempt does not dead-letter healthy messages, and low enough that a poison message does not occupy consumers for long. Values in the single digits to low tens are common. Pair it with a visibility timeout above your worst-case processing time.


Every message in a dead-letter queue is a failure that already happened. Tomosu helps you see the consumer paths that will put the next ones there, before they ship. Assess your repository →