A dead letter queue (DLQ) is a separate queue, or topic, where a messaging system parks messages it could not process: ones that kept failing after a set number of attempts, expired, or were rejected. The main queue keeps moving, and the parked messages stay available to inspect, fix and replay.
This guide explains how a DLQ works, which failures belong in one, how Amazon SQS, Azure Service Bus, RabbitMQ, Apache Kafka and Google Cloud Pub/Sub implement it, and how retries and idempotency keys make a replay safe. Our guides to webhooks and system integration patterns both point here for the failure lane. Vendor details are as documented in October 2026.
How does a dead letter queue work?
A consumer receives a message, processes it and acknowledges it, and the broker deletes it. If the consumer never acknowledges, the broker offers the message again. The broker counts those deliveries, and when a message crosses a limit it moves the message to another queue instead of offering it again. The main queue carries on with the next message. That other queue is the DLQ, which vendors also call a dead-letter exchange, sub-queue or topic (DLT).

The delivery limit is not the only trigger. RabbitMQ lists four: a consumer rejects the message without requeueing it, its time to live (TTL) expires, its queue exceeds a length limit, or it is returned to a quorum queue more times than the delivery limit. Azure Service Bus adds a message whose header exceeds the size quota, and Amazon EventBridge sends an event straight to the DLQ, with no retries, when the target is missing or permission is denied.
Without a DLQ, a message that always fails has nowhere to go. On Amazon SQS it becomes visible again each time its visibility timeout expires, and it keeps coming back until the queue's retention period, 4 days by default, runs out and it is deleted. A DLQ is not an error log: it holds the whole message, so a person or a job can examine it and send it through again. Nor does it replace retries. Retries come first, and the DLQ takes what they could not save.
Which failures should be retried, and which parked?
Retrying is worth it only if the same message could succeed later, and platforms already encode that split. Amazon SNS retries server-side errors but not client-side ones, such as a deleted endpoint. Spring for Apache Kafka skips retries for deserialization and conversion exceptions, which a retry is unlikely to resolve.
| Kind of failure | Examples | Does a retry help? | What to do |
|---|---|---|---|
| Transient | Timeout, connection reset, HTTP 429 or 503 | Usually | Retry with backoff and jitter, honour Retry-After, dead-letter when the attempts run out |
| Long outage | A partner is down for hours | Not within the retry window | Let the limit trip, raise the alarm, redrive after recovery |
| Bad message | Malformed JSON, a missing field, an unknown product code | No | Dead-letter at once with the reason, fix the data or the mapping, redrive |
| Bug in your code | An exception on one shape of payload | No, until you deploy | Dead-letter, fix, redrive |
| Systemic | Expired credential, revoked permission, wrong endpoint | No, and it hits every message | Pause the consumer and page the owner instead of draining the queue into the DLQ |
A poison message is the extreme case. RabbitMQ describes it as a message that causes a consumer to repeatedly requeue a delivery, so it is never positively acknowledged. Without a delivery limit it stays at the head of the line for good, and ordering makes this worse. On an SQS FIFO queue, a failed message blocks its message group until it is deleted or expires. A Kafka consumer group's position is one offset per partition, so a record that always fails holds up everything behind it in that partition until the application sets it aside.
Some failures are ambiguous. The Amazon Builders' Library notes that eventual consistency blurs the line between client and server errors: a client error one moment can turn into a success the next, as state propagates. A 404 for a record created a second ago is a case in point. Classify by error code, not only by status class. RFC 9110, for one, lets a client repeat a request after a 408 Request Timeout.
How do the main platforms implement dead-lettering?
The table compares the settings that matter. The notes after it cover the traps.
| Platform | Where the DLQ is set | Attempt limit (default) | Getting messages back |
|---|---|---|---|
| Amazon SQS | Redrive policy on the source queue; same account and Region; a FIFO queue needs a FIFO DLQ | maxReceiveCount, 1 to 1,000 in the console (API reference default: 10) | Built-in redrive task |
| Amazon SNS | Redrive policy on a subscription; the target is an SQS queue | Server-side errors retried (up to 100,015 times over 23 days for SQS and Lambda endpoints); client-side errors never | A Lambda function or your own consumer |
| Amazon EventBridge | A DLQ on a rule target; standard SQS queues only | Retry policy: 24 hours and 185 attempts | A Lambda function or your own consumer |
| AWS Lambda (asynchronous) | On-failure destination, or a function-level DLQ | MaximumRetryAttempts 0 to 2 (default 2) | Consume the queue or topic |
| Azure Service Bus | Built-in $deadletterqueue sub-queue on every queue and subscription | Maximum delivery count (10), and TTL expiry | Service Bus Explorer, or receive and resend |
| RabbitMQ | dead-letter-exchange policy key on the queue | Quorum queues: delivery-limit (20 since 4.0) | Consume the DLX's queue and republish |
| Apache Kafka | Not in plain consumer groups; Connect sinks, Streams and Spring add it | Connect: fail fast. Spring: 10 attempts, then a log line | Consume the dead-letter topic and republish |
| Google Cloud Pub/Sub | Dead-letter topic on the subscription | Maximum delivery attempts, 5 to 100 (5) | Subscribe to the dead-letter topic and republish |
Warning
A limit without a destination turns a retry loop into data loss. RabbitMQ drops a message whose delivery count passes the limit when no dead-letter exchange is set, and drops dead-lettered messages silently if the exchange does not exist. Lambda discards an asynchronous event that fails all attempts unless you configure a destination. Create the destination first, then set the limit.
Amazon SQS, SNS, EventBridge and Lambda
SQS does not create the DLQ for you, and a FIFO queue can only use a FIFO DLQ. AWS also advises against a DLQ on a FIFO queue when exact order matters. The policy is a JSON string inside the queue attributes, in the format of the AWS CLI reference; the API reference lists the defaults. Save this as attributes.json and apply it with aws sqs set-queue-attributes --queue-url <url> --attributes file://attributes.json:
{
"RedrivePolicy": "{\"deadLetterTargetArn\":\"arn:aws:sqs:us-east-1:111122223333:order-payments-dlq\",\"maxReceiveCount\":\"5\"}",
"VisibilityTimeout": "60"
}
Standard queues have a monitoring quirk. When maxReceiveCount is above 3, SQS moves a message to the back of the queue after its third receive without a delete, and ApproximateAgeOfOldestMessage leaves poison-pill messages out. An age alarm catches a stalled consumer, not one bad message, so alarm on the DLQ as well.
When Lambda reads from SQS, put the DLQ on the queue, not the function. AWS recommends a maxReceiveCount of at least 5 here, and a queue visibility timeout of at least six times the function timeout. A failed batch returns every message to the queue by default, and each redelivery raises its receive count; partial batch responses (ReportBatchItemFailures) return only the failed IDs. For asynchronous invocations, Lambda retries twice by default and you can tune MaximumRetryAttempts and MaximumEventAgeInSeconds. Prefer an on-failure destination to the function-level DLQ: it supports more targets and its record includes the function's response. Our serverless computing guide covers Lambda's invocation models.
The built-in SQS redrive works only for DLQs whose sources are SQS queues, so a DLQ fed by SNS, EventBridge or Lambda is drained by a Lambda function or a consumer you write.
Azure Service Bus
Every queue and every topic subscription has a built-in dead-letter sub-queue, addressed as <queue path>/$deadletterqueue. Service Bus dead-letters a message when its delivery count passes the maximum (10 by default), when it expires with dead-lettering on expiry enabled, when its header exceeds the size quota, or after more than four auto-forwarding hops. Your code can also dead-letter a message on purpose, with a reason and a description. Two details bite. The delivery count rises when a message is abandoned or its lock expires, so closing a receiver before settling a message eventually dead-letters it with the reason MaxDeliveryCountExceeded. And a DLQ ignores time to live and has no automatic cleanup, so messages stay until someone retrieves and completes them. Service Bus Explorer in the Azure portal peeks at DLQ messages, lets you edit them, and resubmits them singly or in batches.
RabbitMQ
A dead-letter exchange (DLX) is a normal exchange attached to a queue by the dead-letter-exchange policy key. RabbitMQ recommends policies over hardcoded x-arguments, which cannot change without redeploying. Each dead-lettering is recorded in an x-death header, and x-first-death-reason and x-last-death-reason hold the first and latest reasons. Dead-lettering is not safe in a cluster by default, because messages are republished without publisher confirms. Quorum queues add an at-least-once strategy that keeps the message until the target confirms it, at the price of more memory and CPU, possible duplicates and overflow set to reject-publish.
Quorum queues also have a delivery-limit, 20 by default since RabbitMQ 4.0. RabbitMQ 4.3 counts only genuine failures against it, so a consumer that returns messages with nack can loop without limit, and adds an opt-in delayed retry with linear back-off: the delay is the smaller of min_delay × delivery_count and max_delay.
rabbitmqctl set_policy orders-dlx "^orders\." \
'{"delivery-limit": 5, "dead-letter-exchange": "orders.dlx", "dead-letter-strategy": "at-least-once", "overflow": "reject-publish"}' \
--apply-to quorum_queues
Apache Kafka
A consumer group's state is one offset per partition, so the broker keeps no failure count per record, and the consumer group protocol has no dead-letter setting. The frameworks around it do:
- Kafka Connect sink connectors fail fast by default: no retries,
errors.tolerance=noneand no dead-letter topic. Seterrors.deadletterqueue.topic.name(optionally witherrors.deadletterqueue.context.headers.enable=true), and never useerrors.tolerance=all, which skips problem records, without a DLQ topic and logging. Once a retry delay reacheserrors.retry.delay.max.ms(60 seconds), Connect adds jitter. - Kafka Streams has
errors.dead.letter.queue.topic.name(default null), spelled differently from Connect's setting. - Spring for Apache Kafka: its default handler makes 10 attempts and logs the record. Add
DeadLetterPublishingRecoverer, which publishes to<originalTopic>-dlton the same partition, so the dead-letter topic needs at least as many partitions, with headers for the original topic, partition, offset and exception. A backoff longer thanmax.poll.interval.msneeds theContainerPausingBackOffHandler.
Google Cloud Pub/Sub
Pub/Sub sets the dead-letter topic on the subscription. The maximum number of delivery attempts defaults to 5 and ranges from 5 to 100, but it is approximate: Pub/Sub may forward a message after fewer attempts or try a few more. Attempts are counted only when the dead-letter topic is configured with the correct IAM permissions, and you need a subscription on the dead-letter topic to read what it receives. Forwarded messages carry attributes such as CloudPubSubDeadLetterSourceDeliveryCount. The default retry policy is immediate redelivery, so a short outage can use up five attempts quickly; switch to exponential backoff, which defaults to a 10 second minimum and a 600 second maximum.
How should retries back off?
Retries are the first line of defence, and they can do harm. The Amazon Builders' Library calls them "selfish": a retrying client spends more of a struggling server's time to raise its own odds. Layers compound the damage. With three retries at each of five layers, a failing database sees 243 times the load, so AWS retries at one point in the stack for low-cost operations, sets a timeout on every remote call and limits the number of retries.
Wait longer after each failure, up to a cap, and randomize the wait. In AWS's simulation of 100 contending clients, plain exponential backoff still produced clusters of calls, while adding jitter cut the number of calls by more than half and finished sooner. "Full jitter" picks a random wait between zero and the capped exponential value. The AWS SDK reference gives it as delay = random(0, 1) × min(20,000 ms, base_delay × 2^retry), with a 50 ms base for transient errors and 1,000 ms for throttling; that behaviour is opt-in through AWS_NEW_RETRIES_2026=true until it becomes the default.

With a queue, the wait is the visibility timeout. On SQS it defaults to 30 seconds and ChangeMessageVisibility sets it per message, up to 12 hours from the first receive. Ask for the ApproximateReceiveCount attribute when you receive, and it tells the consumer which attempt this is:
import random
class Transient(Exception):
retry_after = 0 # seconds from a Retry-After header, when there is one
class Permanent(Exception):
pass
def backoff(attempt, base=30.0, cap=900.0):
"""Full jitter: a random wait between 0 and the capped exponential ceiling."""
return random.uniform(0, min(cap, base * 2 ** (attempt - 1)))
def handle(sqs, queue_url, dlq_url, message):
receipt = {"QueueUrl": queue_url, "ReceiptHandle": message["ReceiptHandle"]}
try:
process(message) # idempotent; see the next section
sqs.delete_message(**receipt)
except Permanent as error: # retrying cannot help: park it now, with the reason
sqs.send_message(QueueUrl=dlq_url, MessageBody=message["Body"], MessageAttributes={
"FailureReason": {"DataType": "String", "StringValue": str(error)[:256]}})
sqs.delete_message(**receipt)
except Transient as error: # leave it on the queue and delay the next attempt
attempt = int(message["Attributes"]["ApproximateReceiveCount"])
wait = max(error.retry_after, backoff(attempt))
sqs.change_message_visibility(**receipt, VisibilityTimeout=min(int(wait), 43200))
Honour the partner's own instruction when it gives one. RFC 9110 defines Retry-After as a number of seconds or an HTTP date, which a 503 response may carry, and RFC 6585 defines 429 Too Many Requests, whose response may include it too. Treat the header as a floor, as the code does, and keep your jitter on top so recovered clients do not arrive together.
A retry budget stops retries from multiplying an outage. The AWS SDK keeps a token bucket of 500 tokens, charges 14 for a transient retry and 5 for a throttling retry, refills it on success and stops retrying when it is empty. The Builders' Library uses a local limit like this instead of circuit breakers, which add behaviour that is hard to test. In a queue consumer, stop retrying in process, fail the message, and let the attempt limit and the DLQ take over. Platforms differ in what they do for you: EventBridge retries for 24 hours with exponential backoff and jitter, while Pub/Sub redelivers at once unless you choose otherwise.
How do idempotency keys make retries and replays safe?
Every platform above retries, and a retry after a lost acknowledgement delivers the same message again, so a consumer must expect duplicates. SQS does not guarantee a message is not delivered more than once, Kafka guarantees at-least-once delivery by default, Lambda can receive the same event more than once, and RabbitMQ's at-least-once dead-lettering can create duplicates in the target queue. A replay from a DLQ is a deliberate duplicate. The goal is not fewer duplicates but harmless ones: handling a message twice has the same effect as handling it once.
Key the consumer on an ID the producer assigns once, such as the order number plus the action, and carry it in the message body or attributes. Do not use the broker's message ID. When SQS redrives a message, it is a new message with a new messageID and enqueue time.
Record the key in your own database, in the same transaction as the local writes. A unique constraint does the checking, as PostgreSQL's ON CONFLICT DO NOTHING shows:
CREATE TABLE processed_messages (
message_key text PRIMARY KEY,
processed_at timestamptz NOT NULL DEFAULT now()
);
BEGIN;
INSERT INTO processed_messages (message_key)
VALUES ('order-1042:charge')
ON CONFLICT DO NOTHING
RETURNING message_key;
-- No row returned: a duplicate. COMMIT, then acknowledge the message.
-- One row returned: run the local writes here, then COMMIT.
COMMIT;
If the insert returns no row, the key already exists, so acknowledge the message and stop. If two copies arrive at the same moment, the unique index lets one insert win, and a crash before the commit rolls the key back together with the writes.

Calls to other systems cannot join your transaction, so pass the key on. RFC 9110 says a client should not automatically retry a request with a non-idempotent method, such as POST, unless it has some means to know the request is idempotent, and an idempotency key is that means. Stripe accepts an Idempotency-Key header on POST requests and saves the status code and body of the first request for each key, including 500 errors, then returns them for later requests with the same key. Keys can be up to 255 characters, and Stripe can remove a key once it is at least 24 hours old. A reused key with different parameters is an error, and a request that failed validation or collided with one still running saves no result, so it can be retried. Rules differ by provider, so read the documentation of each API you call.
Tip
The provider's key window is shorter than your DLQ's. A DLQ message can wait days (14 at most on SQS), and Stripe can drop a key after 24 hours. For a late replay, rely on your own table, and look up the existing object by your order ID before creating a new one. Stripe's metadata field exists to attach your IDs to its objects for lookups.
The IETF's httpapi group has a draft, The Idempotency-Key HTTP Header Field, that would standardize the header. It is an Internet-Draft, not a standard: its latest revision (-07) is dated 15 October 2025, and the Datatracker lists it as expired. If you publish an API of your own, enterprise API development in 2026 covers versioning, rate limits and keys from the provider's side.
On the producing side, saving a record and publishing a message are two writes, and either can fail. AWS Prescriptive Guidance describes the transactional outbox: write the event to an outbox table in the same transaction as the change, and let a separate process publish it. That process can send duplicates, so consumers must still be idempotent. Change data capture is the alternative to an outbox table; our guide to building data analytics software shows how log-based capture works.
Be careful with "exactly once". Kafka's design documentation warns that many systems claim it and that the fine print matters. Kafka supports it for reading, processing and writing between Kafka topics, through Kafka Streams or a transactional producer with a read_committed consumer; for an outside system it needs that system's cooperation. SQS FIFO deduplicates sends within a 5-minute window, which does not cover a replay days later. In practice, exactly once means at-least-once delivery plus an idempotent consumer.
How do you operate a dead letter queue?
Alarm on depth first
Alarm when the DLQ holds one message. On SQS, alarm on ApproximateNumberOfMessagesVisible of the DLQ with a threshold of 1, as the SNS documentation also advises. Do not use NumberOfMessagesSent, which does not count messages that SQS moves to a DLQ itself. Add an alarm on the source queue's ApproximateAgeOfOldestMessage for stalled consumers. Failing to reach the DLQ has its own signal: EventBridge reports InvocationsFailedToBeSentToDLQ and Lambda emits DeadLetterErrors. Pub/Sub records subscription/dead_letter_message_count for forwarded messages.
Keep the reason, the retention and the privacy rules
Store why each message failed. Service Bus has DeadLetterReason and DeadLetterErrorDescription, RabbitMQ the x-death headers, EventBridge ERROR_CODE and ERROR_MESSAGE, and Lambda the request ID and the first 1 KB of the error message. With an SQS redrive policy the cause lives in your logs, so log the message ID on every failed attempt.
Retention is a trap. A standard SQS queue keeps a message's original enqueue timestamp when it moves it, so a message that spent a day in the source queue is deleted 3 days later from a DLQ with 4-day retention. Set the DLQ's retention longer than the source's. The maximum is 14 days, which the SNS documentation recommends for its DLQs.
A DLQ is a second copy of your data with a long retention. Give it the source's encryption and access rules (a redrive that touches an encrypted SQS queue needs kms:Decrypt on its key), and keep secrets and personal data out of failure reasons and logs: Kafka Connect itself warns that logging record contents may log sensitive information. Our secure coding checklist covers logging events rather than sensitive data.
Triage, fix, redrive
- Look before you move anything. Peek at a few messages and their reasons, and read the logs for the exceptions.
- Group by cause: outage, bad data, code bug or systemic failure. Each has a different fix and owner.
- Fix the cause and deploy, or wait until the partner is healthy.
- Redrive one message first, then the rest at a capped rate while you watch the source queue and the partner's error rate.
- Verify and record. The DLQ is empty, reconciliation matches and nothing happened twice. If a class of error should never have been retried, add it to the non-retryable list.
SQS has a built-in redrive task. By default it moves messages, oldest first, from the DLQ to their source queue, or to another queue of the same type. AWS recommends starting with a small rate and ramping up. The maximum is 500 messages per second, a task runs for at most 36 hours, and SQS cannot filter or modify messages during a redrive.
aws sqs start-message-move-task \
--source-arn arn:aws:sqs:us-east-1:111122223333:order-payments-dlq \
--max-number-of-messages-per-second 5
Because the task cannot filter, replaying only some messages means a small script: receive from the DLQ, check the reason, send the ones you want to the source queue, and delete them from the DLQ only after the send succeeds. A crash between the two steps then creates a duplicate rather than a loss, which an idempotent consumer absorbs. On Kafka, consume the dead-letter topic and produce the records back to the original topic.
A DLQ ages badly. Low-grade noise builds up until people mute the alarm, and an old message meets a consumer that no longer understands its shape. Give each DLQ a named owner on the team that owns the consumer, and rehearse a redrive before you need one.
A worked example: an order sync that survives a partner outage
This example is illustrative, not a client. A store publishes an order.paid message to an SQS queue, order-payments. A worker charges the customer's stored payment method through a payment provider and then books the order in an ERP. The queue has a 60-second visibility timeout, maxReceiveCount of 5, a DLQ with 14-day retention and an alarm at one visible DLQ message. The worker records the key in its own table, then sends Idempotency-Key: order-1042:charge to the provider. All times and counts are made up.
| When | What happens | What protects the customer |
|---|---|---|
| Fri 17:58 | The worker receives order 1042. The provider's API times out after 10 seconds, so the charge may or may not have happened. | The key is stored and sent on every attempt |
| 17:58 to 18:15 | Attempts 2 to 5 wait up to 30, 60, 120 and 240 seconds (full jitter) and time out. Where the provider answers 503 with Retry-After, the worker waits at least that long. | Backoff spreads the load on a failing partner |
| About 18:15 | 212 messages have used their five receives and moved to the DLQ. The alarm pages the on-call engineer, who sees timeouts to one host, confirms the provider's incident and does not redrive. | The orders wait safely, for up to 14 days |
| Fri 21:10 | The provider recovers. One message is redriven and succeeds, and the rest follow at 5 per second. Order 1042's first request had reached the provider, which returns its saved result for the same key. | The idempotency key prevents the second charge |
| 21:25 to Mon 09:30 | Three messages fail with a validation error (a missing postal code) and go straight back to the DLQ with the reason. Support corrects the records on Monday and they are redriven. | A permanent error is not retried; triage, then replay |
Now suppose the redrive had happened more than 24 hours later. Stripe, for one, can drop a key after that, so a provider like it might charge order 1042 again. The defence is the worker's own table plus a lookup: before creating a charge, check your table for a completed row for the order, and ask the provider whether a charge carrying order 1042 in its metadata already exists. That is why the idempotency record belongs to you.
Which mistakes make a dead letter queue useless?
| Mistake | Why it hurts | Instead |
|---|---|---|
| Retrying forever | The queue never drains and the partner gets constant load. RabbitMQ calls the no-limit behaviour of 3.13 not recommended, and Spring's FixedBackOff.UNLIMITED_ATTEMPTS and Connect's errors.retry.timeout of -1 retry without end | Set an attempt limit and a maximum age, and let them trip into the DLQ |
| A DLQ nobody watches | Messages expire after the retention period (4 days by default on SQS) and are gone | Alarm at one message and give the DLQ an owner |
| Replaying without idempotency | Every replay risks a duplicate charge, email or record | A stored key, passed on to partner APIs |
| One DLQ for unrelated flows | Owners and causes blur, and SQS cannot filter a redrive, so you replay everything or nothing | One DLQ per source queue or flow |
| A visibility timeout shorter than the work | The message reappears while the first consumer holds it, so it is processed twice and its receive count climbs towards the DLQ | Set the timeout above the processing time, or extend it with a heartbeat |
When is a dead letter queue the wrong tool?
- Strict order per key. A DLQ lets later messages overtake a failed one, which AWS warns can break exact order. If message 7 must be applied before message 8, halt that key or partition, alert, fix and resume instead of skipping.
- A log you can rewind. A Kafka consumer can rewind to an old offset and re-consume, and a lakehouse keeps the raw copy of every record (see medallion architecture). Use a DLQ for individual records that must not block the stream, and a rewind for a bad deploy that affected thousands.
- A systemic failure or a waiting caller. If every message fails, the problem is not the messages: pause the consumer and let the queue hold the backlog. For a synchronous request, return the error to the caller.
A DLQ also costs something to run: another queue to monitor, a retention decision, access rules and a runbook. A flow whose failures you would never replay may need only an alarm and a log.
Add a dead letter queue to one flow this week
- Pick one asynchronous flow with a real cost when it fails, such as orders to accounting, and name the team that owns its consumer.
- Create the DLQ before the limit: same type as the source queue, retention longer than the source's, matching encryption.
- Set the attempt limit and the backoff together, so the total backoff covers the outages you expect.
- Classify errors in the consumer. Transient failures raise a retryable error; bad messages and known bugs are dead-lettered at once with a reason.
- Make the consumer idempotent: a producer-assigned key, a unique constraint, and the same key sent to every partner API.
- Alarm and route. An alarm at one DLQ message and an age alarm on the source queue, both routed to the owning team.
- Rehearse a redrive in a test environment with a deliberately failing message, and write down the steps.
Computese builds this failure lane as part of its integrations service: dead-letter queues, alerts, safe replay, idempotency keys and exponential backoff with jitter, with a runbook for each flow. If your integrations duplicate or lose records after outages, that is the place to start.


