Queues and background jobs
Some work doesn’t belong in the request. Sending an email, resizing an image or calling a slow partner API can take seconds and can fail for reasons unrelated to the user’s action. If you do it inline, the user waits, and a failure in the side task fails the whole request. Moving it to a background job lets you respond quickly and retry the slow part on its own.
Queues and workers
The request handler puts a message describing the work onto a queue and returns. Separate worker processes take messages off the queue and do the work.
// In the request handler
await queue.send("send-welcome-email", { userId: user.id });
return res.status(201).json(user);
// In a worker process
queue.process("send-welcome-email", async (job) => {
const user = await db.users.find(job.data.userId);
await mailer.sendWelcome(user);
});
This decouples the two sides. You can scale workers independently, absorb bursts by letting the queue grow, and keep serving requests while a downstream service is down. Pass IDs rather than whole objects in messages, so the worker reads current data when it runs.
At-least-once delivery
Most queues promise at-least-once delivery: every message is delivered, but some may be delivered more than once. A worker can finish the work and crash before acknowledging it, and the queue will hand the message to someone else. Exactly-once delivery across a network is not something you can rely on.
So handlers must be idempotent: running them twice has the same effect as running them once. Common techniques:
- Record processed message IDs in a table with a unique constraint, and skip duplicates.
- Make writes conditional, such as
UPDATE ... WHERE status = 'pending'. - Pass an idempotency key to external APIs, such as payment providers, that support one.
Visibility timeouts
When a worker takes a message, many queues hide it from other workers for a visibility timeout rather than deleting it. If the worker acknowledges in time, the message is removed. If not, it becomes visible again and another worker picks it up. Set the timeout comfortably longer than the job normally takes, or extend it while a long job runs; otherwise a slow but healthy job gets processed twice in parallel.
Retries and dead-letter queues
Jobs fail. Retry them, but with exponential backoff: wait longer after each failure, so a struggling dependency gets room to recover. Cap the number of attempts.
After the last attempt, move the message to a dead-letter queue (DLQ) instead of dropping it or retrying forever. A message that can never succeed, such as one with malformed data, would otherwise block or burn resources indefinitely. Alert on DLQ growth, inspect the messages, fix the cause, and replay them.
Scheduling
Recurring work, such as nightly reports or cleanup, is usually triggered by a scheduler that enqueues a job at set times. If you run several instances, make sure only one enqueues each run, or make the job idempotent so a duplicate run is harmless. Missed runs during a deploy or outage should be caught up or explicitly skipped, not silently lost.
The outbox pattern
A common bug: you commit an order to the database, then publish an OrderPlaced event, and the process crashes in between. The order exists but nobody is told. Publishing first has the opposite problem.
The outbox pattern fixes this. In the same database transaction as the business change, insert the event into an outbox table. A separate relay reads unsent rows, publishes them, and marks them sent.
BEGIN;
INSERT INTO orders (id, total) VALUES (9001, 4200);
INSERT INTO outbox (topic, payload)
VALUES ('order-placed', '{"orderId": 9001}');
COMMIT;
Because both rows commit atomically, the event is published if and only if the order exists. The relay may publish an event twice if it crashes after sending, so consumers still need to be idempotent.
Habits
- Respond fast; enqueue anything slow or unreliable.
- Assume every message can arrive twice, and write handlers accordingly.
- Retry with backoff and a cap, then send to a dead-letter queue.
- Monitor queue depth, age of the oldest message and DLQ size.
- Use an outbox when a database write and an event must happen together.