Observability for web applications: logs, metrics and traces that earn their cost
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
In the first fifteen minutes of an incident: declare it, assign one person to coordinate, communicate status to users, and mitigate — roll back, disable a flag, scale up — before diagnosing. Root cause analysis belongs in the postmortem, not in the outage.
| Severity | Definition | Response |
|---|---|---|
| SEV1 | Complete outage or data loss risk | All hands, page immediately, 30-min updates |
| SEV2 | Major feature broken, or severe degradation | Page on-call, fix during business hours if scoped |
| SEV3 | Minor feature broken with a workaround | Ticket, next working day |
| SEV4 | Cosmetic or low impact | Backlog |
Write these down before you need them. Arguing about whether something is a SEV1 while it is happening wastes the exact minutes that matter most.
In a team of four, one person may hold several of these — but the roles should be named explicitly. The failure mode without them is three people changing configuration at once and nobody knowing which change did what.
Track two numbers over time: mean time to detect and mean time to recover. Reducing them is worth more than any single prevention measure, because the next incident will have a cause you did not anticipate.
If you promise availability, someone must be reachable. Keep it fair and paid, keep pages actionable, and make sure the rotation has enough people that it is sustainable.
As soon as customer impact is confirmed. Users discovering an outage themselves and finding no acknowledgement is worse for trust than the outage.
Action items with owners and dates, reviewed in your normal planning. Postmortems whose actions are never scheduled are documentation of a repeat incident.
Document what you ruled out, add the instrumentation whose absence blocked you, and close it. Unexplained incidents are common; unexplained incidents with no new observability are a choice.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
Zero downtime is not a deployment tool setting. It is a property of an application that can run two versions at once.
You do not have backups. You have restores — and you only know whether you have those if you have done one this quarter.
No spam. Just the occasional case study and craft breakdown.