# Dead letter queue: how DLQs work, with retries, backoff and idempotency

> A dead letter queue holds messages that failed after set attempts so you can inspect and replay them. How SQS, Service Bus, RabbitMQ, Kafka and Pub/Sub differ.

- URL: https://computese.com/dead-letter-queue/
- Author: Duong Quan Nguyen, CEO, Computese
- Published: 2026-10-04
- Updated: 2026-10-09
- Topics: Integrations, Cloud

## In short
- A dead letter queue is a separate queue or topic where a broker parks messages that failed after a set number of attempts, expired or were rejected, so one bad message cannot block or loop through the main queue.
- Retry only transient failures, with exponential backoff, full jitter, a cap and a retry budget. Send permanent failures (bad payload, schema change, bug) to the DLQ at once, because retrying cannot fix them.
- Defaults differ: Service Bus allows 10 deliveries, RabbitMQ quorum queues 20, Pub/Sub 5 (range 5 to 100), and SQS is set with maxReceiveCount. Plain Kafka consumer groups have no built-in DLQ.
- Delivery is at least once, so duplicates are normal. Make consumers idempotent with a key stored in your own database; a provider's key window (Stripe may drop a key after 24 hours) will not cover a replay days later.
- A DLQ nobody watches is data loss with extra steps. Alarm when it holds one message, keep its retention longer than the source's, and redrive in small, rate-limited batches after the fix.

A dead letter queue (DLQ) is a separate queue, or topic, where a messaging system parks messages it could not process: ones that kept failing after a set number of attempts, expired, or were rejected. The main queue keeps moving, and the parked messages stay available to inspect, fix and replay.

This guide explains how a DLQ works, which failures belong in one, how Amazon SQS, Azure Service Bus, RabbitMQ, Apache Kafka and Google Cloud Pub/Sub implement it, and how retries and idempotency keys make a replay safe. Our guides to [webhooks](https://computese.com/what-is-a-webhook/) and [system integration patterns](https://computese.com/system-integration/) both point here for the failure lane. Vendor details are as documented in October 2026.

## How does a dead letter queue work?

A consumer receives a message, processes it and acknowledges it, and the broker deletes it. If the consumer never acknowledges, the broker offers the message again. The broker counts those deliveries, and when a message crosses a limit it moves the message to another queue instead of offering it again. The main queue carries on with the next message. That other queue is the DLQ, which vendors also call a dead-letter exchange, sub-queue or topic (DLT).

![Envelopes ride a conveyor belt toward a server. One torn, crooked envelope is pushed by an arm onto an orange side shelf below the belt while the rest keep moving.](https://computese.com/images/blog/dead-letter-queue/side-shelf.ec05045cfd-1536.webp)

*The failing message leaves the main line, so everything behind it keeps moving.*

The delivery limit is not the only trigger. [RabbitMQ lists four](https://www.rabbitmq.com/docs/dlx): a consumer rejects the message without requeueing it, its time to live (TTL) expires, its queue exceeds a length limit, or it is returned to a quorum queue more times than the delivery limit. [Azure Service Bus](https://learn.microsoft.com/en-us/azure/service-bus-messaging/service-bus-dead-letter-queues) adds a message whose header exceeds the size quota, and [Amazon EventBridge](https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-rule-dlq.html) sends an event straight to the DLQ, with no retries, when the target is missing or permission is denied.

Without a DLQ, a message that always fails has nowhere to go. On Amazon SQS it [becomes visible again](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html) each time its visibility timeout expires, and it keeps coming back until the queue's retention period, [4 days by default](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/quotas-messages.html), runs out and it is deleted. A DLQ is not an error log: it holds the whole message, so a person or a job can examine it and send it through again. Nor does it replace retries. Retries come first, and the DLQ takes what they could not save.

## Which failures should be retried, and which parked?

Retrying is worth it only if the same message could succeed later, and platforms already encode that split. [Amazon SNS](https://docs.aws.amazon.com/sns/latest/dg/sns-dead-letter-queues.html) retries server-side errors but not client-side ones, such as a deleted endpoint. [Spring for Apache Kafka](https://docs.spring.io/spring-kafka/reference/kafka/annotation-error-handling.html) skips retries for deserialization and conversion exceptions, which a retry is unlikely to resolve.

| Kind of failure  | Examples                                                 | Does a retry help?            | What to do                                                                                 |
| ---------------- | -------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------ |
| Transient        | Timeout, connection reset, HTTP 429 or 503               | Usually                       | Retry with backoff and jitter, honour `Retry-After`, dead-letter when the attempts run out |
| Long outage      | A partner is down for hours                              | Not within the retry window   | Let the limit trip, raise the alarm, redrive after recovery                                |
| Bad message      | Malformed JSON, a missing field, an unknown product code | No                            | Dead-letter at once with the reason, fix the data or the mapping, redrive                  |
| Bug in your code | An exception on one shape of payload                     | No, until you deploy          | Dead-letter, fix, redrive                                                                  |
| Systemic         | Expired credential, revoked permission, wrong endpoint   | No, and it hits every message | Pause the consumer and page the owner instead of draining the queue into the DLQ           |

A **poison message** is the extreme case. [RabbitMQ describes it](https://www.rabbitmq.com/docs/quorum-queues) as a message that causes a consumer to repeatedly requeue a delivery, so it is never positively acknowledged. Without a delivery limit it stays at the head of the line for good, and ordering makes this worse. On an SQS FIFO queue, [a failed message blocks its message group](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-available-cloudwatch-metrics.html) until it is deleted or expires. A Kafka consumer group's position is [one offset per partition](https://kafka.apache.org/43/design/design/), so a record that always fails holds up everything behind it in that partition until the application sets it aside.

Some failures are ambiguous. The [Amazon Builders' Library](https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter) notes that eventual consistency blurs the line between client and server errors: a client error one moment can turn into a success the next, as state propagates. A 404 for a record created a second ago is a case in point. Classify by error code, not only by status class. [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110.html), for one, lets a client repeat a request after a 408 Request Timeout.

## How do the main platforms implement dead-lettering?

The table compares the settings that matter. The notes after it cover the traps.

| Platform                  | Where the DLQ is set                                                                       | Attempt limit (default)                                                                                              | Getting messages back                            |
| ------------------------- | ------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ |
| Amazon SQS                | Redrive policy on the source queue; same account and Region; a FIFO queue needs a FIFO DLQ | `maxReceiveCount`, 1 to 1,000 in the console (API reference default: 10)                                             | Built-in redrive task                            |
| Amazon SNS                | Redrive policy on a subscription; the target is an SQS queue                               | Server-side errors retried (up to 100,015 times over 23 days for SQS and Lambda endpoints); client-side errors never | A Lambda function or your own consumer           |
| Amazon EventBridge        | A DLQ on a rule target; standard SQS queues only                                           | Retry policy: 24 hours and 185 attempts                                                                              | A Lambda function or your own consumer           |
| AWS Lambda (asynchronous) | On-failure destination, or a function-level DLQ                                            | `MaximumRetryAttempts` 0 to 2 (default 2)                                                                            | Consume the queue or topic                       |
| Azure Service Bus         | Built-in `$deadletterqueue` sub-queue on every queue and subscription                      | Maximum delivery count (10), and TTL expiry                                                                          | Service Bus Explorer, or receive and resend      |
| RabbitMQ                  | `dead-letter-exchange` policy key on the queue                                             | Quorum queues: `delivery-limit` (20 since 4.0)                                                                       | Consume the DLX's queue and republish            |
| Apache Kafka              | Not in plain consumer groups; Connect sinks, Streams and Spring add it                     | Connect: fail fast. Spring: 10 attempts, then a log line                                                             | Consume the dead-letter topic and republish      |
| Google Cloud Pub/Sub      | Dead-letter topic on the subscription                                                      | Maximum delivery attempts, 5 to 100 (5)                                                                              | Subscribe to the dead-letter topic and republish |

> [!WARNING]
> A limit without a destination turns a retry loop into data loss. RabbitMQ [drops a message](https://www.rabbitmq.com/docs/quorum-queues) whose delivery count passes the limit when no dead-letter exchange is set, and drops dead-lettered messages silently if the exchange does not exist. Lambda [discards an asynchronous event](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-error-handling.html) that fails all attempts unless you configure a destination. Create the destination first, then set the limit.

### Amazon SQS, SNS, EventBridge and Lambda

SQS does not create the DLQ for you, and a FIFO queue [can only use a FIFO DLQ](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-dead-letter-queue.html). AWS also [advises against a DLQ on a FIFO queue](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html) when exact order matters. The policy is a JSON string inside the queue attributes, in the format of the [AWS CLI reference](https://docs.aws.amazon.com/cli/latest/reference/sqs/set-queue-attributes.html); the [API reference](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/APIReference/API_SetQueueAttributes.html) lists the defaults. Save this as `attributes.json` and apply it with `aws sqs set-queue-attributes --queue-url <url> --attributes file://attributes.json`:

```json
{
  "RedrivePolicy": "{\"deadLetterTargetArn\":\"arn:aws:sqs:us-east-1:111122223333:order-payments-dlq\",\"maxReceiveCount\":\"5\"}",
  "VisibilityTimeout": "60"
}
```

Standard queues have a monitoring quirk. When `maxReceiveCount` is above 3, SQS moves a message to the back of the queue after its third receive without a delete, and `ApproximateAgeOfOldestMessage` [leaves poison-pill messages out](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-available-cloudwatch-metrics.html). An age alarm catches a stalled consumer, not one bad message, so alarm on the DLQ as well.

When Lambda reads from SQS, put the DLQ [on the queue, not the function](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-retain-records.html). AWS [recommends a `maxReceiveCount` of at least 5](https://docs.aws.amazon.com/lambda/latest/dg/services-sqs-configure.html) here, and a queue visibility timeout of at least six times the function timeout. A failed batch returns every message to the queue by default, and each redelivery raises its receive count; [partial batch responses](https://docs.aws.amazon.com/lambda/latest/dg/services-sqs-errorhandling.html) (`ReportBatchItemFailures`) return only the failed IDs. For asynchronous invocations, Lambda [retries twice by default](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-error-handling.html) and you can tune [`MaximumRetryAttempts` and `MaximumEventAgeInSeconds`](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-configuring.html). Prefer an on-failure destination to the function-level DLQ: it supports more targets and its record includes the function's response. Our [serverless computing guide](https://computese.com/the-future-of-serverless-computing-streamlining/) covers Lambda's invocation models.

The built-in SQS redrive works only for DLQs whose sources are SQS queues, so a DLQ fed by SNS, EventBridge or Lambda is drained by a Lambda function or a consumer you write.

### Azure Service Bus

Every queue and every topic subscription has a built-in dead-letter sub-queue, addressed as `<queue path>/$deadletterqueue`. Service Bus dead-letters a message when its delivery count passes the maximum (10 by default), when it expires with dead-lettering on expiry enabled, when its header exceeds the size quota, or after more than four auto-forwarding hops. Your code can also dead-letter a message on purpose, with a reason and a description. Two details bite. The delivery count rises when a message is abandoned or its lock expires, so closing a receiver before settling a message eventually dead-letters it with the reason `MaxDeliveryCountExceeded`. And a DLQ ignores time to live and has no automatic cleanup, so messages stay until someone retrieves and completes them. Service Bus Explorer in the Azure portal peeks at DLQ messages, lets you edit them, and resubmits them singly or in batches.

### RabbitMQ

A dead-letter exchange (DLX) is a normal exchange attached to a queue by the `dead-letter-exchange` policy key. RabbitMQ [recommends policies over hardcoded `x-arguments`](https://www.rabbitmq.com/docs/dlx), which cannot change without redeploying. Each dead-lettering is recorded in an `x-death` header, and `x-first-death-reason` and `x-last-death-reason` hold the first and latest reasons. Dead-lettering is not safe in a cluster by default, because messages are republished without publisher confirms. Quorum queues add an `at-least-once` strategy that keeps the message until the target confirms it, at the price of more memory and CPU, possible duplicates and `overflow` set to `reject-publish`.

Quorum queues also have a `delivery-limit`, 20 by default since RabbitMQ 4.0. RabbitMQ 4.3 counts only genuine failures against it, so a consumer that returns messages with `nack` can loop without limit, and adds an opt-in [delayed retry](https://www.rabbitmq.com/docs/quorum-queues) with linear back-off: the delay is the smaller of `min_delay × delivery_count` and `max_delay`.

```bash
rabbitmqctl set_policy orders-dlx "^orders\." \
  '{"delivery-limit": 5, "dead-letter-exchange": "orders.dlx", "dead-letter-strategy": "at-least-once", "overflow": "reject-publish"}' \
  --apply-to quorum_queues
```

### Apache Kafka

A consumer group's state is one offset per partition, so the broker keeps no failure count per record, and the consumer group protocol has no dead-letter setting. The frameworks around it do:

- **Kafka Connect sink connectors** [fail fast by default](https://kafka.apache.org/43/kafka-connect/user-guide/): no retries, `errors.tolerance=none` and no dead-letter topic. Set `errors.deadletterqueue.topic.name` (optionally with `errors.deadletterqueue.context.headers.enable=true`), and never use `errors.tolerance=all`, which skips problem records, without a DLQ topic and logging. Once a retry delay reaches `errors.retry.delay.max.ms` (60 seconds), [Connect adds jitter](https://kafka.apache.org/43/configuration/kafka-connect-configs/).
- **Kafka Streams** has [`errors.dead.letter.queue.topic.name`](https://kafka.apache.org/43/configuration/kafka-streams-configs/) (default null), spelled differently from Connect's setting.
- **Spring for Apache Kafka**: its default handler makes 10 attempts and logs the record. Add `DeadLetterPublishingRecoverer`, which publishes to `<originalTopic>-dlt` on the same partition, so the dead-letter topic needs at least as many partitions, with headers for the original topic, partition, offset and exception. A backoff longer than `max.poll.interval.ms` needs the `ContainerPausingBackOffHandler`.

### Google Cloud Pub/Sub

Pub/Sub sets the dead-letter topic on the [subscription](https://docs.cloud.google.com/pubsub/docs/dead-letter-topics). The maximum number of delivery attempts defaults to 5 and ranges from 5 to 100, but it is approximate: Pub/Sub may forward a message after fewer attempts or try a few more. Attempts are counted only when the dead-letter topic is configured with the correct IAM permissions, and you need a subscription on the dead-letter topic to read what it receives. Forwarded messages carry attributes such as `CloudPubSubDeadLetterSourceDeliveryCount`. The default [retry policy](https://docs.cloud.google.com/pubsub/docs/subscription-retry-policy) is immediate redelivery, so a short outage can use up five attempts quickly; switch to exponential backoff, which defaults to a 10 second minimum and a 600 second maximum.

## How should retries back off?

Retries are the first line of defence, and they can do harm. The [Amazon Builders' Library](https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter) calls them "selfish": a retrying client spends more of a struggling server's time to raise its own odds. Layers compound the damage. With three retries at each of five layers, a failing database sees 243 times the load, so AWS retries at one point in the stack for low-cost operations, sets a timeout on every remote call and limits the number of retries.

Wait longer after each failure, up to a cap, and randomize the wait. In [AWS's simulation](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/) of 100 contending clients, plain exponential backoff still produced clusters of calls, while adding jitter cut the number of calls by more than half and finished sooner. "Full jitter" picks a random wait between zero and the capped exponential value. The [AWS SDK reference](https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html) gives it as `delay = random(0, 1) × min(20,000 ms, base_delay × 2^retry)`, with a 50 ms base for transient errors and 1,000 ms for throttling; that behaviour is opt-in through `AWS_NEW_RETRIES_2026=true` until it becomes the default.

![Two laptops each send five envelope arrows toward a server along a timeline. The gaps between attempts grow wider, the two rows never line up, and the longest gap in the top row is orange.](https://computese.com/images/blog/dead-letter-queue/backoff-gaps.68314c07f4-1536.webp)

*Growing gaps protect the partner; randomizing them stops every client returning at the same instant.*

With a queue, the wait is the visibility timeout. On SQS it [defaults to 30 seconds](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html) and `ChangeMessageVisibility` sets it per message, up to 12 hours from the first receive. Ask for the [`ApproximateReceiveCount`](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/APIReference/API_ReceiveMessage.html) attribute when you receive, and it tells the consumer which attempt this is:

```python
import random


class Transient(Exception):
    retry_after = 0  # seconds from a Retry-After header, when there is one


class Permanent(Exception):
    pass


def backoff(attempt, base=30.0, cap=900.0):
    """Full jitter: a random wait between 0 and the capped exponential ceiling."""
    return random.uniform(0, min(cap, base * 2 ** (attempt - 1)))


def handle(sqs, queue_url, dlq_url, message):
    receipt = {"QueueUrl": queue_url, "ReceiptHandle": message["ReceiptHandle"]}
    try:
        process(message)  # idempotent; see the next section
        sqs.delete_message(**receipt)
    except Permanent as error:  # retrying cannot help: park it now, with the reason
        sqs.send_message(QueueUrl=dlq_url, MessageBody=message["Body"], MessageAttributes={
            "FailureReason": {"DataType": "String", "StringValue": str(error)[:256]}})
        sqs.delete_message(**receipt)
    except Transient as error:  # leave it on the queue and delay the next attempt
        attempt = int(message["Attributes"]["ApproximateReceiveCount"])
        wait = max(error.retry_after, backoff(attempt))
        sqs.change_message_visibility(**receipt, VisibilityTimeout=min(int(wait), 43200))
```

Honour the partner's own instruction when it gives one. [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110.html) defines `Retry-After` as a number of seconds or an HTTP date, which a 503 response may carry, and [RFC 6585](https://www.rfc-editor.org/rfc/rfc6585.html) defines 429 Too Many Requests, whose response may include it too. Treat the header as a floor, as the code does, and keep your jitter on top so recovered clients do not arrive together.

A retry budget stops retries from multiplying an outage. The AWS SDK keeps a token bucket of 500 tokens, charges 14 for a transient retry and 5 for a throttling retry, refills it on success and stops retrying when it is empty. The Builders' Library uses a local limit like this instead of circuit breakers, which add behaviour that is hard to test. In a queue consumer, stop retrying in process, fail the message, and let the attempt limit and the DLQ take over. Platforms differ in what they do for you: EventBridge retries for 24 hours [with exponential backoff and jitter](https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-rule-retry-policy.html), while Pub/Sub redelivers at once unless you choose otherwise.

## How do idempotency keys make retries and replays safe?

Every platform above retries, and a retry after a lost acknowledgement delivers the same message again, so a consumer must expect duplicates. SQS does not guarantee a message is not [delivered more than once](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html), Kafka [guarantees at-least-once delivery by default](https://kafka.apache.org/43/design/design/), Lambda can [receive the same event more than once](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-error-handling.html), and RabbitMQ's at-least-once dead-lettering can create duplicates in the target queue. A replay from a DLQ is a deliberate duplicate. The goal is not fewer duplicates but harmless ones: handling a message twice has the same effect as handling it once.

Key the consumer on an ID the producer assigns once, such as the order number plus the action, and carry it in the message body or attributes. Do not use the broker's message ID. When SQS redrives a message, [it is a new message](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-dead-letter-queue-redrive.html) with a new `messageID` and enqueue time.

Record the key in your own database, in the same transaction as the local writes. A unique constraint does the checking, as [PostgreSQL's `ON CONFLICT DO NOTHING`](https://www.postgresql.org/docs/current/sql-insert.html) shows:

```sql
CREATE TABLE processed_messages (
  message_key  text PRIMARY KEY,
  processed_at timestamptz NOT NULL DEFAULT now()
);

BEGIN;
INSERT INTO processed_messages (message_key)
VALUES ('order-1042:charge')
ON CONFLICT DO NOTHING
RETURNING message_key;
-- No row returned: a duplicate. COMMIT, then acknowledge the message.
-- One row returned: run the local writes here, then COMMIT.
COMMIT;
```

If the insert returns no row, the key already exists, so acknowledge the message and stop. If two copies arrive at the same moment, the unique index lets one insert win, and a crash before the commit rolls the key back together with the writes.

![Two identical envelopes with the same key tag travel from a cloud toward a gear-shaped worker. The first is recorded on a ledger with an orange check mark and reaches the database; the second is turned back.](https://computese.com/images/blog/dead-letter-queue/key-stamped-once.594c83b9ff-1536.webp)

*Two copies arrive, one effect happens: the ledger remembers the key.*

Calls to other systems cannot join your transaction, so pass the key on. [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110.html) says a client should not automatically retry a request with a non-idempotent method, such as POST, unless it has some means to know the request is idempotent, and an idempotency key is that means. [Stripe](https://docs.stripe.com/api/idempotent_requests) accepts an `Idempotency-Key` header on POST requests and saves the status code and body of the first request for each key, including 500 errors, then returns them for later requests with the same key. Keys can be up to 255 characters, and Stripe can remove a key once it is at least 24 hours old. A reused key with different parameters is an error, and a request that failed validation or collided with one still running saves no result, so it can be retried. Rules differ by provider, so read the documentation of each API you call.

> [!TIP]
> The provider's key window is shorter than your DLQ's. A DLQ message can wait days (14 at most on SQS), and Stripe can drop a key after 24 hours. For a late replay, rely on your own table, and look up the existing object by your order ID before creating a new one. Stripe's `metadata` field exists to attach your IDs to its objects for lookups.

The IETF's httpapi group has a draft, [The Idempotency-Key HTTP Header Field](https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/), that would standardize the header. It is an Internet-Draft, not a standard: its latest revision (-07) is dated 15 October 2025, and the Datatracker lists it as expired. If you publish an API of your own, [enterprise API development in 2026](https://computese.com/the-future-of-enterprise-api-development/) covers versioning, rate limits and keys from the provider's side.

On the producing side, saving a record and publishing a message are two writes, and either can fail. AWS Prescriptive Guidance describes the [transactional outbox](https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html): write the event to an outbox table in the same transaction as the change, and let a separate process publish it. That process can send duplicates, so consumers must still be idempotent. Change data capture is the alternative to an outbox table; our guide to [building data analytics software](https://computese.com/building-data-analytics-software/) shows how log-based capture works.

Be careful with "exactly once". Kafka's design documentation warns that many systems claim it and that [the fine print matters](https://kafka.apache.org/43/design/design/). Kafka supports it for reading, processing and writing between Kafka topics, through Kafka Streams or a transactional producer with a `read_committed` consumer; for an outside system it needs that system's cooperation. SQS FIFO deduplicates sends within a [5-minute window](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-available-cloudwatch-metrics.html), which does not cover a replay days later. In practice, exactly once means at-least-once delivery plus an idempotent consumer.

## How do you operate a dead letter queue?

### Alarm on depth first

Alarm when the DLQ holds one message. On SQS, [alarm on `ApproximateNumberOfMessagesVisible`](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/dead-letter-queues-alarms-cloudwatch.html) of the DLQ with a threshold of 1, as the SNS documentation also advises. Do not use `NumberOfMessagesSent`, which does not count messages that SQS moves to a DLQ itself. Add an alarm on the source queue's `ApproximateAgeOfOldestMessage` for stalled consumers. Failing to reach the DLQ has its own signal: EventBridge reports `InvocationsFailedToBeSentToDLQ` and Lambda emits `DeadLetterErrors`. Pub/Sub records `subscription/dead_letter_message_count` for forwarded messages.

### Keep the reason, the retention and the privacy rules

Store why each message failed. Service Bus has `DeadLetterReason` and `DeadLetterErrorDescription`, RabbitMQ the `x-death` headers, EventBridge `ERROR_CODE` and `ERROR_MESSAGE`, and Lambda the request ID and the first 1 KB of the error message. With an SQS redrive policy the cause lives in your logs, so log the message ID on every failed attempt.

Retention is a trap. A standard SQS queue keeps a message's original enqueue timestamp when it moves it, so [a message that spent a day in the source queue is deleted 3 days later from a DLQ with 4-day retention](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html). Set the DLQ's retention longer than the source's. The maximum is 14 days, which the SNS documentation recommends for its DLQs.

A DLQ is a second copy of your data with a long retention. Give it the source's encryption and access rules (a [redrive](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-dead-letter-queue-redrive.html) that touches an encrypted SQS queue needs `kms:Decrypt` on its key), and keep secrets and personal data out of failure reasons and logs: Kafka Connect itself warns that logging record contents [may log sensitive information](https://kafka.apache.org/43/kafka-connect/user-guide/). Our [secure coding checklist](https://computese.com/best-practices-for-secure-coding/) covers logging events rather than sensitive data.

### Triage, fix, redrive

1. **Look before you move anything.** Peek at a few messages and their reasons, and read the logs for the exceptions.
2. **Group by cause:** outage, bad data, code bug or systemic failure. Each has a different fix and owner.
3. **Fix the cause and deploy,** or wait until the partner is healthy.
4. **Redrive one message first,** then the rest at a capped rate while you watch the source queue and the partner's error rate.
5. **Verify and record.** The DLQ is empty, reconciliation matches and nothing happened twice. If a class of error should never have been retried, add it to the non-retryable list.

SQS has a [built-in redrive task](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-dead-letter-queue-redrive.html). By default it moves messages, oldest first, from the DLQ to their source queue, or to another queue of the same type. AWS recommends starting with a small rate and ramping up. The maximum is [500 messages per second](https://docs.aws.amazon.com/cli/latest/reference/sqs/start-message-move-task.html), a task runs for at most 36 hours, and SQS cannot filter or modify messages during a redrive.

```bash
aws sqs start-message-move-task \
  --source-arn arn:aws:sqs:us-east-1:111122223333:order-payments-dlq \
  --max-number-of-messages-per-second 5
```

Because the task cannot filter, replaying only some messages means a small script: receive from the DLQ, check the reason, send the ones you want to the source queue, and delete them from the DLQ only after the send succeeds. A crash between the two steps then creates a duplicate rather than a loss, which an idempotent consumer absorbs. On Kafka, consume the dead-letter topic and produce the records back to the original topic.

A DLQ ages badly. Low-grade noise builds up until people mute the alarm, and an old message meets a consumer that no longer understands its shape. Give each DLQ a named owner on the team that owns the consumer, and rehearse a redrive before you need one.

## A worked example: an order sync that survives a partner outage

This example is illustrative, not a client. A store publishes an `order.paid` message to an SQS queue, `order-payments`. A worker charges the customer's stored payment method through a payment provider and then books the order in an ERP. The queue has a 60-second visibility timeout, `maxReceiveCount` of 5, a DLQ with 14-day retention and an alarm at one visible DLQ message. The worker records the key in its own table, then sends `Idempotency-Key: order-1042:charge` to the provider. All times and counts are made up.

| When               | What happens                                                                                                                                                                                            | What protects the customer                            |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------- |
| Fri 17:58          | The worker receives order 1042. The provider's API times out after 10 seconds, so the charge may or may not have happened.                                                                              | The key is stored and sent on every attempt           |
| 17:58 to 18:15     | Attempts 2 to 5 wait up to 30, 60, 120 and 240 seconds (full jitter) and time out. Where the provider answers 503 with `Retry-After`, the worker waits at least that long.                              | Backoff spreads the load on a failing partner         |
| About 18:15        | 212 messages have used their five receives and moved to the DLQ. The alarm pages the on-call engineer, who sees timeouts to one host, confirms the provider's incident and does not redrive.            | The orders wait safely, for up to 14 days             |
| Fri 21:10          | The provider recovers. One message is redriven and succeeds, and the rest follow at 5 per second. Order 1042's first request had reached the provider, which returns its saved result for the same key. | The idempotency key prevents the second charge        |
| 21:25 to Mon 09:30 | Three messages fail with a validation error (a missing postal code) and go straight back to the DLQ with the reason. Support corrects the records on Monday and they are redriven.                      | A permanent error is not retried; triage, then replay |

Now suppose the redrive had happened more than 24 hours later. Stripe, for one, can drop a key after that, so a provider like it might charge order 1042 again. The defence is the worker's own table plus a lookup: before creating a charge, check your table for a completed row for the order, and ask the provider whether a charge carrying order 1042 in its metadata already exists. That is why the idempotency record belongs to you.

## Which mistakes make a dead letter queue useless?

| Mistake                                    | Why it hurts                                                                                                                                                                                                                          | Instead                                                                  |
| ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| Retrying forever                           | The queue never drains and the partner gets constant load. RabbitMQ calls the no-limit behaviour of 3.13 not recommended, and Spring's `FixedBackOff.UNLIMITED_ATTEMPTS` and Connect's `errors.retry.timeout` of -1 retry without end | Set an attempt limit and a maximum age, and let them trip into the DLQ   |
| A DLQ nobody watches                       | Messages expire after the retention period (4 days by default on SQS) and are gone                                                                                                                                                    | Alarm at one message and give the DLQ an owner                           |
| Replaying without idempotency              | Every replay risks a duplicate charge, email or record                                                                                                                                                                                | A stored key, passed on to partner APIs                                  |
| One DLQ for unrelated flows                | Owners and causes blur, and SQS cannot filter a redrive, so you replay everything or nothing                                                                                                                                          | One DLQ per source queue or flow                                         |
| A visibility timeout shorter than the work | The message reappears while the first consumer holds it, so it is processed twice and its receive count climbs towards the DLQ                                                                                                        | Set the timeout above the processing time, or extend it with a heartbeat |

## When is a dead letter queue the wrong tool?

- **Strict order per key.** A DLQ lets later messages overtake a failed one, which AWS warns can break exact order. If message 7 must be applied before message 8, halt that key or partition, alert, fix and resume instead of skipping.
- **A log you can rewind.** A Kafka consumer can [rewind to an old offset and re-consume](https://kafka.apache.org/43/design/design/), and a lakehouse keeps the raw copy of every record (see [medallion architecture](https://computese.com/medallion-architecture/)). Use a DLQ for individual records that must not block the stream, and a rewind for a bad deploy that affected thousands.
- **A systemic failure or a waiting caller.** If every message fails, the problem is not the messages: pause the consumer and let the queue hold the backlog. For a synchronous request, return the error to the caller.

A DLQ also costs something to run: another queue to monitor, a retention decision, access rules and a runbook. A flow whose failures you would never replay may need only an alarm and a log.

## Add a dead letter queue to one flow this week

1. **Pick one asynchronous flow** with a real cost when it fails, such as orders to accounting, and name the team that owns its consumer.
2. **Create the DLQ before the limit:** same type as the source queue, retention longer than the source's, matching encryption.
3. **Set the attempt limit and the backoff together,** so the total backoff covers the outages you expect.
4. **Classify errors in the consumer.** Transient failures raise a retryable error; bad messages and known bugs are dead-lettered at once with a reason.
5. **Make the consumer idempotent:** a producer-assigned key, a unique constraint, and the same key sent to every partner API.
6. **Alarm and route.** An alarm at one DLQ message and an age alarm on the source queue, both routed to the owning team.
7. **Rehearse a redrive** in a test environment with a deliberately failing message, and write down the steps.

Computese builds this failure lane as part of its [integrations service](https://computese.com/services/integrations/): dead-letter queues, alerts, safe replay, idempotency keys and exponential backoff with jitter, with a runbook for each flow. If your integrations duplicate or lose records after outages, that is the place to start.

## Key terms
- **Dead letter queue (DLQ)**: A separate queue or topic that receives messages a broker could not process or deliver, so they can be inspected and replayed. Also called a dead-letter topic (DLT) or, in RabbitMQ, a dead-letter exchange (DLX).
- **Poison message**: A message that makes a consumer fail every time it is delivered, so it is never acknowledged. Without a delivery limit it is redelivered forever and can block the messages behind it.
- **Redrive**: Moving messages from a DLQ back to a queue for processing. Amazon SQS has a built-in redrive task with a rate limit; on most other platforms you republish the messages yourself.
- **Visibility timeout**: On Amazon SQS, the time a received message stays hidden from other consumers. If the consumer does not delete it in time, it becomes visible again and counts as another delivery attempt.
- **Delivery count**: The number of times a broker has handed a message to a consumer. SQS calls it the receive count; Service Bus and RabbitMQ call it the delivery count. Exceeding a configured limit sends the message to the DLQ.
- **Exponential backoff with jitter**: A retry schedule in which the wait grows after each failure, up to a cap, and is randomized so that many clients do not retry at the same moment. Full jitter picks a random wait between zero and the capped value.
- **Retry budget**: A cap on the extra traffic that retries may add, such as a token bucket that retries spend and successes refill, so retries cannot multiply an outage.
- **Idempotency key**: A unique value sent with a request so the receiver can recognize a repeat and return the first result instead of acting twice. Stripe's Idempotency-Key header is a widely used example.
- **At-least-once delivery**: A guarantee that a message is not lost but may be delivered more than once. SQS standard queues and Kafka (by default) work this way, so consumers must tolerate duplicates.
- **Transactional outbox**: A pattern in which an event is written to a table in the same database transaction as the data change, and a separate process publishes it, so the change and its message cannot disagree.

## Common questions

### What is a dead letter queue?

A dead letter queue is a separate queue or topic where a messaging system puts messages it could not process: they failed too many times, expired, were rejected or could not be delivered. It keeps one bad message from blocking or endlessly looping through the main queue, and it keeps the message so you can inspect it, fix the cause and send it through again.

### How many times should a message be retried before it goes to the DLQ?

There is no universal number. It depends on how long transient failures last and how far apart the attempts are. As of October 2026, Service Bus defaults to 10 deliveries, RabbitMQ quorum queues to 20, Pub/Sub to 5 (range 5 to 100) and the SQS console accepts 1 to 1,000, with AWS advising at least 5 for Lambda consumers. Pick the limit together with your backoff schedule so the total retry time covers the outages you expect.

### Does Apache Kafka have a dead letter queue?

Not in the consumer group protocol. Kafka Connect sink connectors can write failed records to a dead-letter topic, Kafka Streams has an errors.dead.letter.queue.topic.name setting, and Spring for Apache Kafka publishes to a topic ending in -dlt with DeadLetterPublishingRecoverer. For any other consumer you build the pattern yourself: after the retries, publish the record and its error to a separate topic, then commit the offset.

### How do I reprocess messages from a dead letter queue?

Fix the cause first, then move the messages back in small, rate-limited batches. SQS has a built-in redrive task (StartMessageMoveTask) capped at 500 messages per second, Service Bus Explorer can resubmit messages, and elsewhere you consume the DLQ and republish. The consumer must be idempotent, because a replay is a duplicate delivery.

### What is a poison message?

A message that makes the consumer fail every time it is delivered, for example a malformed payload or a bug triggered by one payload shape. Retrying cannot help, and without a delivery limit the message is redelivered forever and can block the messages behind it. A limit plus a DLQ removes it from the flow and keeps it for diagnosis.

### What is an idempotency key, and how long does it last?

It is a unique value a client sends with a request so the server can recognize a retry and return the first result instead of acting twice. Stripe accepts keys of up to 255 characters on POST requests and can remove them once they are at least 24 hours old; other providers differ, and the IETF header draft is not a standard. Keep your own record of processed keys for replays that arrive later.

### How long should a dead letter queue keep messages?

Longer than the source queue keeps them, and long enough for someone to notice and act. On SQS a message moved from a standard queue keeps its original enqueue time, so a DLQ with the same retention as the source can expire a message almost at once. The SQS maximum is 14 days, which AWS recommends for SNS dead-letter queues.

## Sources
1. [Using dead-letter queues in Amazon SQS](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html), Amazon Web Services
2. [Configure a dead-letter queue using the Amazon SQS console](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-dead-letter-queue.html), Amazon Web Services
3. [Learn how to configure a dead-letter queue redrive in Amazon SQS](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-configure-dead-letter-queue-redrive.html), Amazon Web Services
4. [start-message-move-task (AWS CLI command reference)](https://docs.aws.amazon.com/cli/latest/reference/sqs/start-message-move-task.html), Amazon Web Services
5. [set-queue-attributes (AWS CLI command reference)](https://docs.aws.amazon.com/cli/latest/reference/sqs/set-queue-attributes.html), Amazon Web Services
6. [SetQueueAttributes (Amazon SQS API Reference)](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/APIReference/API_SetQueueAttributes.html), Amazon Web Services
7. [ReceiveMessage (Amazon SQS API Reference)](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/APIReference/API_ReceiveMessage.html), Amazon Web Services
8. [Amazon SQS visibility timeout](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html), Amazon Web Services
9. [Amazon SQS message quotas](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/quotas-messages.html), Amazon Web Services
10. [Available CloudWatch metrics for Amazon SQS](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-available-cloudwatch-metrics.html), Amazon Web Services
11. [Creating alarms for dead-letter queues using Amazon CloudWatch](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/dead-letter-queues-alarms-cloudwatch.html), Amazon Web Services
12. [Amazon SNS dead-letter queues](https://docs.aws.amazon.com/sns/latest/dg/sns-dead-letter-queues.html), Amazon Web Services
13. [Using dead-letter queues to process undelivered events in EventBridge](https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-rule-dlq.html), Amazon Web Services
14. [How EventBridge retries delivering events](https://docs.aws.amazon.com/eventbridge/latest/userguide/eb-rule-retry-policy.html), Amazon Web Services
15. [How Lambda handles errors and retries with asynchronous invocation](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-error-handling.html), Amazon Web Services
16. [Configuring error handling settings for Lambda asynchronous invocations](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-configuring.html), Amazon Web Services
17. [Capturing records of Lambda asynchronous invocations](https://docs.aws.amazon.com/lambda/latest/dg/invocation-async-retain-records.html), Amazon Web Services
18. [Creating and configuring an Amazon SQS event source mapping](https://docs.aws.amazon.com/lambda/latest/dg/services-sqs-configure.html), Amazon Web Services
19. [Handling errors for an SQS event source in Lambda](https://docs.aws.amazon.com/lambda/latest/dg/services-sqs-errorhandling.html), Amazon Web Services
20. [Service Bus dead-letter queues](https://learn.microsoft.com/en-us/azure/service-bus-messaging/service-bus-dead-letter-queues), Microsoft Learn
21. [Dead Letter Exchanges (RabbitMQ 4.3)](https://www.rabbitmq.com/docs/dlx), RabbitMQ
22. [Quorum Queues (RabbitMQ 4.3)](https://www.rabbitmq.com/docs/quorum-queues), RabbitMQ
23. [Design: message delivery semantics (Apache Kafka 4.3)](https://kafka.apache.org/43/design/design/), Apache Kafka
24. [Kafka Connect user guide: error reporting (Apache Kafka 4.3)](https://kafka.apache.org/43/kafka-connect/user-guide/), Apache Kafka
25. [Kafka Connect configs (Apache Kafka 4.3)](https://kafka.apache.org/43/configuration/kafka-connect-configs/), Apache Kafka
26. [Kafka Streams configs (Apache Kafka 4.3)](https://kafka.apache.org/43/configuration/kafka-streams-configs/), Apache Kafka
27. [Handling Exceptions (Spring for Apache Kafka 4.1)](https://docs.spring.io/spring-kafka/reference/kafka/annotation-error-handling.html), Spring
28. [Dead-letter topics](https://docs.cloud.google.com/pubsub/docs/dead-letter-topics), Google Cloud Pub/Sub
29. [Subscription retry policy](https://docs.cloud.google.com/pubsub/docs/subscription-retry-policy), Google Cloud Pub/Sub
30. [Exponential Backoff And Jitter](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/), AWS Architecture Blog
31. [Timeouts, retries, and backoff with jitter](https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter), Amazon Builders' Library
32. [Retry behavior](https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html), AWS SDKs and Tools Reference Guide
33. [RFC 9110: HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110.html), IETF
34. [RFC 6585: Additional HTTP Status Codes](https://www.rfc-editor.org/rfc/rfc6585.html), IETF
35. [Idempotent requests](https://docs.stripe.com/api/idempotent_requests), Stripe API Reference
36. [The Idempotency-Key HTTP Header Field (Internet-Draft)](https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/), IETF Datatracker
37. [Transactional outbox pattern](https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html), AWS Prescriptive Guidance
38. [INSERT (PostgreSQL 18)](https://www.postgresql.org/docs/current/sql-insert.html), PostgreSQL Documentation
