Stabilization playbook
Status: Stable Techniques that made a visual-regression suite reproducible, with the evidence behind each one. Every entry names the consumer it was proven in and the measurement that proved it. Read this before adding a mask, an ignore-list entry, or a wait.
The first question: does the fix change pixels?
Split every candidate stabiliser into two classes before considering it.
A timing stabiliser makes the capture wait longer for the page to reach the state it was always going to reach. Waiting for image decode, waiting for network idle, waiting for fonts. It can only move the capture toward the true render, so it is safe to default on.
A pixel-altering stabiliser makes the page render differently from the way a user sees it. image-rendering: pixelated, hiding an element, swapping an asset. It invalidates every existing baseline and puts a lie in the reference images, so it needs evidence that the instability is real and that no timing fix reaches it.
The library shipped one of each in consecutive candidates and conflated them. rc.6 set image-rendering: pixelated on every image to stop a scaler non-determinism that does not exist in Chromium 1055: six sequential captures of a 1600-to-1425 downscale, six of a 1220-to-740 downscale, six with the decode margin removed and eight concurrent full-page captures were all byte-identical. The property itself changed 367,580 pixels, 28.4% of a capture, at pixelmatch threshold 0. It was reverted in rc.7. See docs/plans/2026-09-19-deterministic-image-scaling.md.
Measure before you attribute
A diff report names a file, not a cause. Three attributions in the EKZ investigation were wrong because they were read off filenames.
Localise the diff spatially first. Compute the bounding box of the differing pixels, then ask the DOM which element occupies that box. On seiten-detailseite--annual-report-page the box was x=194..1037 y=1034..1088, inside ekz-highlight-box-grid — which identified the counter row in one step, after two wrong guesses from the story name.
Correlate across the whole run. EKZ build 17 reported 3,658 failures in 12,288 captures. All 3,627 diff markers sat in image-bearing components (Detailseite 430, Seiteneinstieg 384, Startseite 330, Bild-Slideshow 265, Bild 160); 2,080 captures of text-only components produced zero (Formular-Komponenten 0/1248, Tabelle 0/160, CTA-Button 0/352, Akkordeon 0/224). A defect that partitions that cleanly has one cause, not eleven.
Check your instrument before believing a negative result. A throwaway shadow-DOM walk that stopped at depth 6 reported two distinct renderings where the shipped depth-10 walk found one. A local repro of the scaler claim used a browser that was not the CI browser, because Chromium 1055 does not launch on an arm64 Mac (spawn error -88) — run the experiment inside the CI image.
Compare against what the pipeline actually does. A determinism sweep that omits the config's maskSelectors measures raw rendering, not the suite. The same sweep reported one unstable story with masks omitted and nothing once they were applied.
Reduced motion reaches CSS, not JavaScript
The library sets the preference on every context and kills CSS animation unconditionally, in src/capture/screenshot.ts (preparePageForScreenshot) and at each context creation in src/core/browser-pool.ts and src/cli/executor.ts:
await page.addStyleTag({ content: `*, *::before, *::after {
animation-duration: 0s !important; animation-delay: 0s !important;
transition-duration: 0s !important; transition-delay: 0s !important; }` });
await page.emulateMedia({ reducedMotion: "reduce" });
That covers CSS animations and transitions and nothing else. A counter driven by requestAnimationFrame, an autoplaying video, a drag inertia loop and a SMIL animation inside a background-image SVG all ignore it, because none of them is a CSS animation.
The app has to read the preference for those. BFH does it with one helper and gates the JS at each site:
// bfh-frontend/source/assets/scripts/lib/prefers-reduced-motion.js
export const prefersReducedMotion = () => matchMedia('(prefers-reduced-motion)')?.matches || false;
// c-hero.js
this.playing = this.config.autoplay && !prefersReducedMotion();
startVideo() { if (prefersReducedMotion()) return; }
This is the preferred fix, because it is the accessibility behaviour the preference already asks for: the test runner gets determinism as a side effect of the app being correct, and no baseline records something a user never sees.
Two cases it does not reach. An animation isolated inside a background-image SVG cannot be gated from the host document at all — BFH swapped in a static per-theme SVG showing a frozen frame, deliberately keeping the affordance visible rather than hiding it. And a component that freezes mid-animation rather than completing needs the end state forced, not the animation stopped: use-count-up with isCounting={onScreen} halts wherever it was when the element left the viewport, so three counters on one EKZ page recorded 360/12/40 on one load and 65/2/7 on another — the same ~18% of each animation, frozen at a timing-dependent point. Waiting longer made it worse (4000ms settle gave 3 distinct captures of 4) because extra time only adds variance to when the freeze happened.
Masks
A Playwright mask paints over an element's bounding box. It stabilises content that changes inside a fixed box. It does not stabilise a box whose own size changes, because then the painted rectangle differs between runs and the masks become the diff.
Test the claim before recording it. EKZ's config carried FIXME these masks do not work over .annual-report-number__counter; masking that selector gives one distinct capture in five, in every tenant theme tried. The evidence behind the FIXME concerned a different selector, .annual-report-number-swiss-map__counter, whose box does change.
BFH made the same error in the other direction and recorded the correction (c42ab2a):
Build 397 shows masked regions coming out magenta once the selectors reach the capture, so the option was never ignored — the selector list was being dropped upstream, which looked the same from outside.
When a mask appears not to work, confirm the selectors reach the capture before concluding anything about masking. Masked regions render magenta; if nothing is magenta, the list was dropped.
Layout and viewport
Width selects behaviour, not only type scale. BFH captured at 601px, which sat 0.02px above a CSS bucket edge and produced up to 46% diff against 600px; and at 768px, where SearchResults.applyInitialView() matched mq-xsmall through an indexOf('small') test and forced list view. Widths moved to 390 / 1024 / 1440 / 1920, each well inside its bucket. EKZ documents the margin beside each width for the same reason.
Images without intrinsic size shift the page during scroll. BFH added imgWidth/imgHeight to 121 image objects across 11 stories and rescoped a selector that had left <img> at height: 0 until load; target drift during scroll went from 2,807px to 0.
Per-run state
Clear web storage in an init script, before navigation, not after. BFH's search-results component persists grid-or-list under bfh-bfh-navigation-history and reads it during its own init; with 32 parallel workers the story order varies per run, so the same story rendered grid in one run and list in the next — 52% diff, no code change.
Strip variable text out of error placeholders. Timestamps, retry counters, UUIDs and file positions in a failure image make every identical failure a fresh diff.
Retry
A retry captures the page again and compares the new capture; it never compares the same two images twice, because pixelmatch is deterministic. Timeouts, HTTP errors, capture failures and size mismatches are retried. A completed pixel comparison that differs is retried only where the project sets retries.onDiff, because BFH measured the cost of retrying every difference at 49 deterministic diffs times two retries times about 12 seconds. EKZ's Seiten/Startseite "Interactive Experience Startseite" differs on a different set of tenants in every run while two Storybook builds produce byte-identical captures of it, so the capture is the variable there and only a second capture settles it. The loop stops as soon as two attempts write the same bytes, so a deterministic difference costs one extra capture. A test that passes after a retry is reported as flaky, with its URL, so the fix starts from the report.
Artifact footprint
Archiving is part of stability, because a full controller fails builds that captured perfectly. EKZ build 18 completed all 12,288 captures and then failed with No space left on device during archiveArtifacts: a run archives about 9GB onto a 40GB volume shared with every other job on the instance. Bound retention per job, and prefer serving images from somewhere other than the controller.