Designing a CI/CD pipeline that people trust
A pipeline that takes 25 minutes and fails randomly does not improve quality. It teaches the team to merge on red.
Test business logic exhaustively with fast unit tests, test components at the interaction level with Testing Library queries that mirror how users find elements, and reserve end-to-end tests for the handful of flows that generate revenue. Coverage percentage is a poor target; coverage of critical paths is the real goal.
| Layer | Test with | How much |
|---|---|---|
| Pure logic (pricing, permissions, formatting) | Unit tests | Thorough, including edge cases |
| Components with interaction | Testing Library | Main states and one failure path |
| Forms and validation | Component tests | Valid, invalid, submitting, error |
| Critical user journeys | Playwright / Cypress | 3–8 flows total |
| Static presentational markup | — | Usually not worth it |
The distribution matters more than the total. A suite of 400 snapshot tests over presentational components will not catch the bug where the discount calculation rounds the wrong way.
Query the way a user would — by role, label and text — and assert on visible outcomes rather than internal state.
test("shows a validation error for an invalid email", async () => {
const user = userEvent.setup();
render(<EnquiryForm />);
await user.type(screen.getByLabelText(/work email/i), "not-an-email");
await user.click(screen.getByRole("button", { name: /send/i }));
expect(await screen.findByRole("alert")).toHaveTextContent(/valid email/i);
});A useful side effect: tests written with role and label queries fail when the markup is inaccessible. If getByRole('button') cannot find your clickable div, your users' assistive technology cannot either.
Fewer than you think. End-to-end tests are the slowest to run, the most expensive to maintain and the most prone to flakiness. Cover the flows where a failure costs money — signup, checkout, the core workflow — and stop.
Coverage measures which lines ran, not whether they were verified. It is useful for finding untested files and useless as a KPI. We track coverage but gate on something better: does every critical path have at least one test, and did the last three production bugs get a regression test?
Sparingly. Large component snapshots break on every legitimate change and get updated without reading. Small snapshots of serialised data structures are more useful than DOM snapshots.
Vitest is generally faster and integrates cleanly with Vite-based tooling; Jest has the larger ecosystem and long-standing Next.js support. Either is fine — pick one and be consistent.
Test the data functions they call as plain units, and cover the rendered output with an end-to-end test. Direct unit testing of async server components has limited tooling support today.
No. The last 20% is usually error paths and glue code where tests cost more than they return. Aim for confidence on the paths that matter rather than a number.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
A pipeline that takes 25 minutes and fails randomly does not improve quality. It teaches the team to merge on red.
Every codebase is well organised on day one. The question is what it looks like after forty feature requests have been bolted onto the same Button.
These are not beginner mistakes. They are the ones that pass review, work in development, and break under real conditions.
No spam. Just the occasional case study and craft breakdown.