Your Background Job Failed Halfway Through… How Do You Safely Retry Without Doing the Work Twice?
How to design retry-safe background jobs with idempotency keys, deduplication, checkpoints, transactional outboxes, leases, and compensation for partial failure.
7 min read
The Job Charged the Card, Then Crashed
An order-completion job performs four steps:
1. Charge payment
2. Mark order paid
3. Reserve inventory
4. Send receiptThe payment succeeds. The database update succeeds. Then the worker crashes before inventory and email.
The queue retries the job from the beginning.
If the handler simply repeats every step, the customer may be charged twice. If it skips the entire job, inventory and email never happen.
The requirement is not “retry the function.” It is:
Resume or replay the workflow without repeating completed effects.
Assume Any Line Can Run More Than Once
Most durable queues provide at-least-once delivery because a worker can finish work and crash before acknowledging the message.
queue delivers job
|
v
worker performs effect
|
X crashes before ACK
|
v
visibility timeout expires
|
v
queue delivers job againThis is the same boundary that appears when a Kafka consumer processes a message before committing its offset.
Design the handler under the assumption that delivery can repeat.
Give the Job a Stable Identity
Bad:
job ID generated again for every retryBetter:
job ID: complete-order:ord_4201Retries of the same logical work must carry the same identity.
await queue.add(
"complete-order",
{ orderId: "ord_4201" },
{ jobId: "complete-order:ord_4201" },
);Queue-level deduplication helps prevent duplicate enqueues. It is not enough by itself because deduplication windows expire and workers can still redeliver an active job.
Make Each Side Effect Idempotent
Database State Transitions
Use a guarded transition:
UPDATE orders
SET payment_status = 'paid'
WHERE id = $1
AND payment_status = 'pending'
RETURNING id;If the order is already paid, a retry does not apply the transition again.
External Payment
Send a stable idempotency key to the provider:
await paymentProvider.charge({
amount: order.total,
idempotencyKey: `order:${order.id}:capture`,
});The provider should return the original result for repeated keys. Store its operation ID locally.
An email provider may accept a message and then lose the response. Use a stable message ID where supported or record the intended email in an outbox before sending.
Exactly-once email delivery is difficult. Define whether a duplicate receipt is preferable to no receipt, and make the subject/body identify the order clearly.
Track Workflow State Durably
For a multi-step job, store step progress:
CREATE TABLE order_completion (
order_id text PRIMARY KEY,
payment_operation_id text,
inventory_reserved_at timestamptz,
receipt_requested_at timestamptz,
completed_at timestamptz,
last_error text,
updated_at timestamptz NOT NULL DEFAULT now()
);On retry:
payment_operation_id exists? skip charge
inventory_reserved_at exists? skip reservation
receipt_requested_at exists? skip outbox insertion
completed_at exists? return successCheckpoints should represent durable effects, not “about to perform” flags.
Bad sequence:
mark payment complete
crash before chargingThe retry incorrectly skips the charge.
Use the external provider’s idempotency key and record its confirmed result.
Keep Database Changes and Outgoing Intent Together
Suppose the job marks an order paid and must publish OrderPaid.
This has a gap:
UPDATE order -> commit
publish event -> crashUse a transactional outbox:
BEGIN;
UPDATE orders
SET payment_status = 'paid'
WHERE id = $1;
INSERT INTO outbox (id, event_type, aggregate_id, payload)
VALUES ($2, 'OrderPaid', $1, $3)
ON CONFLICT (id) DO NOTHING;
COMMIT;An outbox worker publishes the event and marks it sent. Publishing must still be idempotent because the worker can crash after publish but before its status update.
Use a Lease, Not a Permanent “Processing” Flag
Two workers can receive the same job during failover. A lease gives temporary ownership:
UPDATE jobs
SET
locked_by = $1,
locked_until = now() + interval '30 seconds'
WHERE id = $2
AND (
locked_until IS NULL
OR locked_until < now()
)
RETURNING *;The worker renews the lease while making progress. If it dies, another worker can continue after expiration.
A permanent status = 'processing' flag can strand work forever after a crash.
Leases do not replace idempotency. A paused worker can wake after its lease expires and overlap with a replacement worker. The effects still need guards.
Partial Completion May Need Compensation
Some effects cannot be rolled back atomically.
payment captured
inventory reservation fails permanentlyPossible business decisions:
- Retry inventory for a bounded period
- Source inventory from another location
- Refund or void the payment
- Move the order to manual review
- Inform the user of delayed fulfillment
Compensation is a new forward action, not time travel. A refund can fail and must also be idempotent and observable.
Classify Failures
Transient:
- Network timeout
- Temporary database failover
- Provider
503 - Rate limit with retry guidance
Permanent:
- Invalid email address
- Missing required order data
- Unsupported currency
- Deleted account
Unknown outcome:
- External call timed out after request transmission
- Worker died after an effect but before recording the result
Unknown outcomes are why stable idempotency keys matter. Do not convert uncertainty into duplicate work.
Retry With Backoff and Limits
attempt 1 -> wait with jitter
attempt 2 -> longer wait
attempt 3 -> longer wait
attempt 4 -> dead-letter / manual reviewTrack the oldest job age, not just queue depth. Ten old stuck jobs can be more serious than a thousand new jobs moving normally.
Diagnose Failed and Duplicate Jobs
Log and measure:
job_id
logical_idempotency_key
attempt
current_step
lease_owner
external_operation_id
failure_class
durationUseful metrics:
job_attempts_total
job_retries_total
job_duplicate_suppressed_total
job_failures_total by class
job_oldest_age_seconds
job_lease_expired_total
compensation_totalWithout a stable job ID across attempts, the same logical job looks like several unrelated failures.
Trade-offs
| Technique | Benefit | Cost |
|---|---|---|
| Queue job ID | Prevents duplicate enqueue | Limited to queue retention semantics |
| Database unique key | Strong deduplication | Requires retention and conflict handling |
| Step checkpoints | Resumes long workflows | More workflow state and migrations |
| Transactional outbox | No lost outgoing intent | Publisher and cleanup operations |
| Lease | Recovers abandoned work | Renewal and overlap complexity |
| Compensation | Repairs partial business outcome | Another fallible workflow |
Production Best Practices
- Assign one stable ID to the logical job.
- Make every repeated effect idempotent.
- Guard database transitions with current state.
- Use provider idempotency keys for external writes.
- Store durable checkpoints after confirmed effects.
- Commit state and outgoing intent through an outbox.
- Use expiring leases for worker ownership.
- Bound retries and classify permanent failures.
- Provide dead-letter inspection and replay tooling.
- Design compensation for irreversible partial outcomes.
Conclusion
A safe retry does not restart a multi-step workflow blindly. It re-enters a durable state machine.
stable job identity
+
idempotent effects
+
durable checkpoints
+
lease and bounded retry
+
compensation when necessaryOnce every boundary can be replayed, a worker crash becomes an expected recovery path instead of a duplicate-payment incident.
Retry delays still consume downstream capacity. The retry-storm guide explains how to bound attempts with deadlines, backoff, jitter, and a shared retry budget.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.