Background jobs and queues: getting work off the request path
Every request that sends an email, generates a file or calls a third party is a request that should have returned already.
The Node.js event loop runs your JavaScript on a single thread, processing callbacks in phases — timers, pending callbacks, poll, check, close — while I/O happens on a thread pool or the OS. Any synchronous work you do blocks every other request, which is why CPU-heavy operations belong in worker threads or a separate service.
It repeatedly asks: is there work ready to run? The loop cycles through phases — timers (setTimeout callbacks whose time has come), pending I/O callbacks, poll (waiting for and processing new I/O), check (setImmediate), and close handlers. Between each callback it drains the microtask queue, where resolved promises live.
Your JavaScript runs on one thread throughout. Concurrency comes from the fact that waiting for a database or a file happens elsewhere — the OS or libuv's thread pool — and hands a callback back when it is done.
| Runs on | Examples |
|---|---|
| Main thread (blocking) | Your code, JSON parsing, template rendering, regex, loops |
| libuv thread pool | File system, DNS lookups, some crypto and compression |
| Kernel / OS | Network sockets, timers |
Response times for every endpoint degrade at once, in proportion to the blocking work, even for endpoints that do nothing.
A report we debugged for a client: p50 latency was 40ms and p99 was 3.2 seconds across every route. The cause was one endpoint synchronously generating a PDF. While it ran, all other requests queued behind it — including the health check, which caused the orchestrator to restart healthy pods.
import { monitorEventLoopDelay } from "node:perf_hooks";
const histogram = monitorEventLoopDelay({ resolution: 20 });
histogram.enable();
setInterval(() => {
const p99ms = histogram.percentile(99) / 1e6;
if (p99ms > 100) logger.warn({ p99ms }, "event loop lag high");
histogram.reset();
}, 30_000);Roughly the number of available CPU cores, minus one for the main thread. Beyond that they contend for CPU and add scheduling overhead without increasing throughput.
Yes, when running Node directly on a multi-core machine. In a container orchestrator it is usually simpler to run one process per container and scale the container count instead.
No. Both use the same single-threaded JavaScript execution model. Their runtimes are faster at some operations, but a blocking loop blocks them too.
Under 10ms at p99 for a typical API is comfortable. Sustained values above 100ms mean requests are queueing and users are feeling it.
Harshal Patel
Founder & Lead Engineer, ROVQIX
Harshal leads engineering at ROVQIX, where he has shipped production Next.js, Node.js and PostgreSQL systems for startups, SaaS teams and ecommerce brands. He writes about the trade-offs behind architecture decisions rather than the framework of the week.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Every request that sends an email, generates a file or calls a third party is a request that should have returned already.
The goal of error handling is not to prevent crashes. It is to make sure that when something fails, you can tell what, where and for whom.
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
No spam. Just the occasional case study and craft breakdown.