Observability for web applications: logs, metrics and traces that earn their cost
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
Separate operational errors — a timeout, a validation failure, a missing record — which you handle and recover from, from programmer errors like a null dereference, which should crash the process so a supervisor restarts it clean. Attach context to every error, log structurally with a request id, and never swallow an error into an empty catch block.
| Operational | Programmer | |
|---|---|---|
| Examples | Timeout, 404, invalid input, disk full | undefined is not a function, bad state |
| Expected? | Yes, in normal operation | No — a bug |
| Response | Handle, retry, return an error | Crash and restart |
| User sees | A meaningful message | A generic 500 |
Trying to recover from a programmer error means continuing with a process whose state you no longer understand. Restarting is not a failure of engineering; it is the only honest response to unknown state.
class AppError extends Error {
constructor(
message: string,
readonly code: string,
readonly status: number,
readonly context: Record<string, unknown> = {},
options?: { cause?: unknown }
) {
super(message, options);
this.name = new.target.name;
}
}
throw new AppError("Payment provider timed out", "payment_timeout", 502,
{ orderId, provider: "stripe", attempt: 2 }, { cause: err });A stack trace tells you where. Context tells you which order, which customer, which attempt — the difference between a five-minute fix and an afternoon of guessing.
logger.error({
err, // serialised with stack and cause
requestId: ctx.requestId,
userId: ctx.userId,
route: "POST /orders",
durationMs: 1832,
}, "order creation failed");A stable machine-readable code, a human-readable message that is safe to display, and the request id. Never a stack trace — it leaks file paths, dependency versions and sometimes queries.
{
"error": {
"code": "payment_timeout",
"message": "We could not reach the payment provider. No charge was made.",
"requestId": "req_01HQ8Z3F2K"
}
}No. Catch where you can do something useful — retry, fall back, add context and rethrow. Catching to log and continue usually turns one clear failure into three confusing ones downstream.
Distinguish retryable from permanent, let the queue handle backoff for the former, and dead-letter the latter with full context. Carry the originating request id into the job payload.
For programmer errors, yes — that is what orchestrator restarts are for. What matters is that the crash is logged with enough detail and that restarts are alerted on, so a crash loop cannot hide.
A standard way to chain errors: new Error('msg', { cause: originalError }). It preserves the original stack so you can see the low-level failure behind a high-level message.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
Node is single-threaded for your code. Every millisecond you spend in a synchronous loop is a millisecond nobody else's request is being served.
Every request that sends an email, generates a file or calls a third party is a request that should have returned already.
No spam. Just the occasional case study and craft breakdown.