Background jobs and queues: getting work off the request path
Every request that sends an email, generates a file or calls a third party is a request that should have returned already.
A well-designed webhook system signs every payload with a timestamped HMAC, retries with exponential backoff for a defined window, guarantees at-least-once delivery with an event id consumers can deduplicate on, and gives integrators a delivery log they can inspect and replay themselves.
{
"id": "evt_01HQ8Z3F2K7YQ",
"type": "order.completed",
"created": "2026-04-03T09:12:44Z",
"apiVersion": "2026-01-15",
"sequence": 4821,
"data": { "orderId": "ord_123", "total": 4900, "currency": "INR" }
}Sign the timestamp and raw body with an HMAC secret, send it in a header, and require consumers to verify it in constant time within a short tolerance window.
// Sender
const signed = `${timestamp}.${rawBody}`;
const signature = crypto.createHmac("sha256", secret).update(signed).digest("hex");
headers["X-Signature"] = `t=${timestamp},v1=${signature}`;
// Receiver
const expected = crypto.createHmac("sha256", secret).update(`${t}.${rawBody}`).digest();
const ok = crypto.timingSafeEqual(expected, Buffer.from(v1, "hex"));
if (!ok || Math.abs(Date.now() / 1000 - Number(t)) > 300) return res.status(400).end();| Response | Action |
|---|---|
| 2xx | Delivered, done |
| 408, 429, 5xx | Retry with exponential backoff |
| Other 4xx | Do not retry — the endpoint rejected it permanently |
| Timeout / connection error | Retry |
Do not promise ordered delivery. Retries, parallel workers and network reordering all break it, and a guarantee you cannot keep is worse than none. Include a monotonically increasing sequence number or a resource version so consumers can discard events older than their current state.
No. They should verify the signature, persist the event, return 200 immediately, and process from a queue. Processing inline causes timeouts and unnecessary retries on your side.
24 to 72 hours with exponential backoff covers most transient outages. Beyond that the event is usually stale; disable the endpoint and notify the customer instead.
Webhooks for timeliness and efficiency, polling as a fallback for consumers who cannot host a public endpoint. Offering a list-events API alongside webhooks lets them reconcile after downtime.
Automatically disable after a threshold of consecutive failures, email the account owner, and keep events retrievable through your API so nothing is lost when they fix it.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Every request that sends an email, generates a file or calls a third party is a request that should have returned already.
An API is a product with developers as users. Most of what makes one good is consistency, not cleverness.
The hard part of versioning is not the URL scheme. It is agreeing on what counts as breaking, and then actually retiring the old version.
No spam. Just the occasional case study and craft breakdown.