Dockerfile best practices: caching, reproducibility and image hygiene
Most Dockerfiles are copied from a tutorial and never revisited. Twenty minutes of attention usually halves both build time and image size.
A production Node.js image should use a multi-stage build to keep build tooling out of the runtime layer, install dependencies before copying source so layer caching works, run as a non-root user, and handle SIGTERM for graceful shutdown. Done properly, a typical Next.js image drops from over a gigabyte to around 150–200MB.
# syntax=docker/dockerfile:1
FROM node:22-alpine AS deps
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --omit=dev
FROM node:22-alpine AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build
FROM node:22-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production
RUN addgroup -S app && adduser -S app -G app
COPY --from=deps --chown=app:app /app/node_modules ./node_modules
COPY --from=build --chown=app:app /app/.next/standalone ./
COPY --from=build --chown=app:app /app/.next/static ./.next/static
COPY --from=build --chown=app:app /app/public ./public
USER app
EXPOSE 3000
CMD ["node", "server.js"]For Next.js, the standalone output mode is what makes this small — it traces exactly the files the server needs and leaves the rest behind. Enable it with output: "standalone" in next.config.
Docker caches each layer and invalidates every layer after the first one that changes, so anything that changes often must come last.
Copying your whole source tree before running npm ci means every one-character source change re-installs all dependencies. Copying only package.json and the lockfile first means the install layer is reused until dependencies actually change — turning a three-minute build into twenty seconds.
When an orchestrator scales down or deploys, it sends SIGTERM and waits before sending SIGKILL. A process that ignores SIGTERM has its in-flight requests killed — visible to users as random errors during every deploy.
const server = app.listen(3000);
for (const signal of ["SIGTERM", "SIGINT"]) {
process.on(signal, async () => {
server.close(async () => { // stop accepting new connections
await db.end(); // drain the pool
process.exit(0);
});
setTimeout(() => process.exit(1), 10_000).unref(); // hard limit
});
}| Approach | Typical size | Notes |
|---|---|---|
| node:22 + full source | 1.1–1.4 GB | The accidental default |
| node:22-slim + multi-stage | 300–400 MB | Good, Debian-based |
| node:22-alpine + multi-stage | 150–200 MB | Smallest common choice |
| distroless runtime | 120–180 MB | No shell — hardest to debug |
Alpine uses musl rather than glibc, which occasionally breaks native modules — sharp, canvas, some database drivers. Test your actual dependency set; if something misbehaves, slim is a perfectly good answer at a modest size cost.
Start with Alpine for size. If a native dependency misbehaves under musl, switch to slim — the extra 150MB is not worth hours of debugging obscure binary issues.
Not on a managed platform that builds and runs it for you. You need it for self-hosting, for consistent CI, and for running the same artifact across environments.
Never use ARG or ENV for secrets — they persist in the image history. Use BuildKit secret mounts at build time and inject runtime secrets from the orchestrator.
Node sizes its heap from the machine, not the container limit, on older versions. Set --max-old-space-size to roughly 75% of the container's memory limit and verify with a load test.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
Most Dockerfiles are copied from a tutorial and never revisited. Twenty minutes of attention usually halves both build time and image size.
Container security starts with four lines in a Dockerfile and a scan in CI. Most breaches involve failures at that level, not exotic escapes.
Zero downtime is not a deployment tool setting. It is a property of an application that can run two versions at once.
No spam. Just the occasional case study and craft breakdown.