Zero downtime deployments: rolling, blue-green and canary compared
Zero downtime is not a deployment tool setting. It is a property of an application that can run two versions at once.
A good pipeline gives fast feedback on every pull request — lint, type check and unit tests in a few minutes — then builds one artifact and promotes that same artifact through environments. Slow checks run in parallel or after merge, and every gate that can block a merge must be reliable enough that a red result always means a real problem.
| Stage | Target time | Blocks merge? |
|---|---|---|
| Lint + type check | < 60s | Yes |
| Unit + component tests | < 3 min | Yes |
| Build + image scan | < 5 min | Yes |
| End-to-end on preview | < 10 min | Yes for critical flows |
| Performance / Lighthouse budget | < 5 min | Warn, then enforce |
| Deploy to staging | automatic on merge | — |
| Deploy to production | on approval or tag | — |
Run independent stages in parallel. Lint, type check and unit tests have no reason to be sequential, and parallelising them typically halves the wall-clock time of a pull request check.
Because rebuilding per environment means the artifact you tested is not the artifact you shipped.
Build a single immutable image tagged by commit SHA, push it once, and deploy that exact digest to staging and then production. Configuration comes from the environment at runtime. If staging passed and production fails, you have eliminated 'the build was different' from the investigation.
- name: Build and push
run: |
docker build -t $REGISTRY/app:$GITHUB_SHA .
docker push $REGISTRY/app:$GITHUB_SHA
- name: Deploy to staging
run: ./deploy.sh staging $REGISTRY/app:$GITHUB_SHA
- name: Deploy to production # same digest, different config
if: github.ref == 'refs/heads/main'
run: ./deploy.sh production $REGISTRY/app:$GITHUB_SHAContinuous deployment works well with strong test coverage, feature flags and fast rollback. Without those, an approval step is a reasonable safety measure — just keep it a click, not a meeting.
Quarantine them immediately so they stop blocking merges, then fix them on a deadline. Automatic retries hide the flakiness and let it spread.
Previews validate a change; staging validates the integrated main branch against production-like data and third-party sandboxes. Most teams benefit from both.
The DORA four: deployment frequency, lead time for changes, change failure rate, and time to restore. They describe delivery health better than pipeline duration alone.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Zero downtime is not a deployment tool setting. It is a property of an application that can run two versions at once.
Most Dockerfiles are copied from a tutorial and never revisited. Twenty minutes of attention usually halves both build time and image size.
A test suite nobody trusts is worse than none — it costs time and provides false confidence. Here is what we actually test on client projects.
No spam. Just the occasional case study and craft breakdown.