Running database migrations without taking the site down
The migration that takes your site down is almost never the complicated one. It is the ALTER TABLE that took a lock nobody expected.
A backup strategy starts with two numbers: RPO, how much data you can afford to lose, and RTO, how long you can be down. Those determine your backup frequency and method. For PostgreSQL, continuous WAL archiving gives point-in-time recovery with an RPO of minutes; the strategy is only real once you have performed a timed restore.
RPO is the maximum acceptable data loss measured in time; RTO is the maximum acceptable time to be back online.
| Business type | Typical RPO | Typical RTO | Implies |
|---|---|---|---|
| Marketing site | 24 hours | 8 hours | Nightly snapshots |
| SaaS application | 5 minutes | 1 hour | WAL archiving + hot standby |
| Ecommerce | 1 minute | 15 minutes | Streaming replication + PITR |
| Payments / ledger | ~0 | Minutes | Synchronous replication, multi-region |
These are business decisions with cost attached, not technical preferences. Ask the owner what an hour of lost orders costs, and the backup budget answers itself.
Restore the most recent base backup, then replay archived write-ahead log segments up to a specified timestamp or transaction. This is what lets you recover to 14:22:59 — one second before someone ran a DELETE without a WHERE clause.
# postgresql.conf
archive_mode = on
archive_command = 'pgbackrest --stanza=main archive-push %p'
# Recovery target
recovery_target_time = '2026-04-10 14:22:59+05:30'
recovery_target_action = 'promote'Base backups daily for most systems, with continuous WAL archiving for a low RPO. If you can only afford daily snapshots, be explicit that your RPO is up to 24 hours.
Often yes for hardware failure, but verify retention, whether point-in-time recovery is enabled, whether backups survive account deletion, and how long a restore actually takes. Then test one.
Long enough to catch slow-burn corruption — 30 to 90 days is common — balanced against storage cost and data protection obligations to delete personal data.
Enable versioning and cross-region replication, and consider object lock for immutability. Uploaded files are frequently the least-protected part of an otherwise careful backup strategy.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
The migration that takes your site down is almost never the complicated one. It is the ALTER TABLE that took a lock nobody expected.
The question is not whether a secret will leak. It is whether you will know, and how long it takes to make the leaked one useless.
Observability is not more dashboards. It is being able to answer a question you did not anticipate, without shipping new code.
No spam. Just the occasional case study and craft breakdown.