Retry, Quarantine, or Fix? A Flaky Test Triage Matrix for Playwright CI Pipelines
A green rerun is not a fix: it is a decision you made without looking, and this matrix makes you look.

The build is red. You open the log, see locator.click: Timeout 5000ms exceeded on a checkout test nobody touched, click “Re-run failed jobs”, and it goes green. You merge. Nobody files anything, because nothing is broken anymore. Multiply that by every engineer on the team and every day of the quarter, and you get a suite where retries: 2 quietly carries a growing pile of bugs that only show up as rerun minutes.
Every flaky test forces one of three decisions: retry it and accept the noise, quarantine it so it stops blocking merges, or fix it now. Each is the right call for some flakes and a costly mistake for others. The problem is that most teams make the call implicitly, one rerun at a time, with no data. This post gives you an explicit triage matrix and the three pieces of Playwright code that feed it: a config that turns retries into a detector, a custom reporter that classifies every failed attempt, and a history script that turns those reports into a decision per test.
The problem is well documented. Google’s testing team reported in Flaky Tests at Google and How We Mitigate Them (2016) that about 1.5% of all their test runs reported a flaky result, and that almost 16% of their tests had some level of flakiness. On root causes, Luo et al.’s An Empirical Analysis of Flaky Tests (FSE 2014) classified 201 commits that fixed flaky tests across 51 open-source projects and found async wait to be the largest category, followed by concurrency and test order dependency. In our experience, browser suites follow the same shape: most of the Playwright flakes we trace end at a wait registered too late, a locator read once instead of asserted, or state shared between tests.
Why does a retry make a flaky test pass?
Answer capsule: Playwright discards the failed worker and retries in a fresh process, so leaked state resets and a warmed-up system wins the race it lost.
To decide whether a retry is safe, you need to know what a retry actually changes. When a Playwright test fails, the runner discards the worker process that ran it, along with its browser, and starts a fresh worker for the next attempt. The new worker gets a new workerIndex, re-runs worker-scoped fixtures from scratch, and launches a clean browser. That is a feature: one bad test cannot poison the rest of the shard. It is also why retries are so good at hiding bugs. Anything that depends on leftover state (a user account created by an earlier test in the same worker, a cached auth token, a websocket that never closed) is reset before the second attempt.
The environment changes too. The first attempt warmed up the dev server’s JIT, filled the database connection pool, and pulled assets into the CDN or browser cache. The retry runs against a faster system, so a race that the first attempt lost (a response arriving before its listener was attached, an animation finishing before a click) often goes the other way. Playwright then labels the test flaky: test.outcome() returns "flaky" when an attempt failed and a later one passed. By default a flaky test does not fail the run. Recent Playwright versions add a failOnFlakyTests config option that flips that, which is useful for new tests on a pull request but too blunt for a whole legacy suite.
One gotcha changes the math. In a test.describe.configure({ mode: 'serial' }) group, a failure retries the whole group, so tests that passed get re-run and their attempt counts go up with them. Count failed attempts, not total attempts, or your flake rates will be inflated for every test that happens to share a serial block with a flaky one.
The triage matrix
The matrix routes each flaky test using three signals: how often it fails, what the failure looks like (its signature), and what you lose if it stops gating merges. The thresholds below are starting points from our own pipelines, not industry standards. Tune them to your suite size and CI cost.
| Signal | Decision | Why |
|---|---|---|
| Fails every retry in 30%+ of runs | Fix now (or revert) | This is a regression, not flakiness |
| Any strict mode, detached element, or wrong-value assertion flake | Fix now | The app rendered a different state; a retry hides a real race |
Tagged @critical, flaky at any rate | Fix now, never quarantine | Losing checkout or login coverage costs more than a red build |
| Flaky in 5%+ of runs, timeout or network signature | Quarantine with owner and expiry | Frequent enough to train people to ignore red builds |
| Flaky in under 5% of runs, network or infra signature | Retry is acceptable | Rare, and the cause sits outside the code under test |
| Fewer than 10 runs of history | Collect more data | One bad night looks like a 10% flake rate |
Signatures do the most work here. A timeout tells you something was slow, which can be infrastructure. A strict mode violation tells you the DOM had two matching elements at that moment and one at another, which is always a property of your app or your locator. Treating both as “flaky, retry it” is how real bugs ship.
Step 1: make retries a detector, not a cure
The config below keeps retries on in CI, but every other line exists to turn a retry into evidence. Quarantine is expressed in code with a tag plus an annotation, so it goes through review like any other change, and the main job filters quarantined tests out with grepInvert. Playwright matches grep against a string built from the project name, file name, describe titles, test title, and tags, so a tag is the most reliable handle.
// playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
const isCI = !!process.env.CI;
// QUARANTINE_MODE=1 runs ONLY quarantined tests, in a separate non-blocking CI job.
const quarantineMode = process.env.QUARANTINE_MODE === '1';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: isCI,
// Retries are a detector, not a cure. Two extra attempts are enough to label
// a test "flaky" without letting a consistently broken test burn CI minutes.
retries: isCI ? 2 : 0,
use: {
baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
// Keep the trace of the attempt that FAILED. 'on-first-retry' records the
// retry instead, which for a flaky test is usually the run that passed.
trace: 'retain-on-first-failure',
screenshot: 'only-on-failure',
},
// The main job never sees quarantined tests; the quarantine job sees nothing else.
...(quarantineMode ? { grep: /@quarantine/ } : { grepInvert: /@quarantine/ }),
reporter: [
['list'],
['json', { outputFile: 'test-results/results.json' }],
[
'./reporters/flaky-triage-reporter.ts',
{ outputFile: 'test-results/triage.json', quarantineJob: quarantineMode },
],
],
projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }],
});Note the trace mode. The common trace: 'on-first-retry' records the first retry, and for a flaky test that is usually the attempt that passed. You get a perfect trace of the run with no bug in it. retain-on-first-failure records the first attempt and keeps it only if it failed, which is the evidence you actually need.
Step 2: classify every failed attempt with a custom reporter
Playwright’s built-in JSON reporter has the raw data, but it is shaped for display, not for decisions. A small custom reporter can emit exactly what the matrix needs: a stable id, the final outcome, and a failure signature per failed attempt. It also enforces quarantine expiry, so a “temporary” quarantine cannot quietly become permanent.
// reporters/flaky-triage-reporter.ts
import * as fs from 'node:fs';
import * as path from 'node:path';
import type { FullResult, Reporter, TestCase, TestResult } from '@playwright/test/reporter';
type Signature =
| 'strict-mode'
| 'detached-element'
| 'network'
| 'action-timeout'
| 'test-timeout'
| 'assertion'
| 'unknown';
export interface TriageEntry {
id: string; // project › file › describe › title
tags: string[];
outcome: 'expected' | 'unexpected' | 'flaky' | 'skipped';
attempts: number;
signatures: Signature[]; // one per genuinely failed attempt
quarantine?: { owner?: string; ticket?: string; expires?: string; expired: boolean };
}
// Order matters: the first match wins, and a strict-mode violation inside a
// click also mentions the action name, so the specific patterns go first.
const SIGNATURES: Array<[Signature, RegExp]> = [
['strict-mode', /strict mode violation/i],
['detached-element', /not attached to the DOM|element was detached/i],
['network', /net::ERR_|ECONNRESET|ECONNREFUSED|socket hang up/i],
['test-timeout', /Test timeout of \d+ms exceeded/i],
['action-timeout', /Timeout \d+ms exceeded/i],
['assertion', /expect\(/i],
];
// Playwright error messages are colourised; strip ANSI codes before matching.
const ANSI = /\u001b\[[0-9;]*m/g;
function classify(result: TestResult): Signature {
const text = result.errors.map((e) => e.message ?? e.value ?? '').join('\n').replace(ANSI, '');
for (const [sig, re] of SIGNATURES) if (re.test(text)) return sig;
return result.status === 'timedOut' ? 'test-timeout' : 'unknown';
}
function parseQuarantine(test: TestCase, now: number): TriageEntry['quarantine'] {
const note = test.annotations.find((a) => a.type === 'quarantine');
if (!note) return undefined;
const fields = Object.fromEntries(
(note.description ?? '')
.split(';')
.map((pair) => pair.split('=').map((s) => s.trim()))
.filter((kv): kv is [string, string] => kv.length === 2 && kv[0] !== ''),
);
const expiresAt = Date.parse(fields.expires ?? '');
// Fail closed: a quarantine without a parseable expiry counts as expired,
// otherwise "temporary" quarantines silently become permanent.
const expired = Number.isNaN(expiresAt) || expiresAt < now;
return { owner: fields.owner, ticket: fields.ticket, expires: fields.expires, expired };
}
export default class FlakyTriageReporter implements Reporter {
private readonly outputFile: string;
private readonly quarantineJob: boolean;
private readonly results = new Map<TestCase, TestResult[]>();
constructor(options: { outputFile?: string; quarantineJob?: boolean } = {}) {
this.outputFile = options.outputFile ?? 'test-results/triage.json';
this.quarantineJob = options.quarantineJob ?? false;
}
// Called once per ATTEMPT, so a test retried twice arrives here three times.
onTestEnd(test: TestCase, result: TestResult) {
const list = this.results.get(test) ?? [];
list.push(result);
this.results.set(test, list);
}
async onEnd(result: FullResult): Promise<{ status?: FullResult['status'] } | void> {
const now = Date.now();
const entries: TriageEntry[] = [];
for (const [test, attempts] of this.results) {
entries.push({
id: test.titlePath().filter(Boolean).join(' › '),
tags: test.tags,
// outcome() already accounts for test.fail(): an expected failure is
// "expected", so we never misreport it as flaky or broken.
outcome: test.outcome(),
attempts: attempts.length,
signatures: attempts
// Evidence = an attempt that did not end as expected. This drops
// test.fail() runs that failed on purpose, and 'interrupted' attempts
// (maxFailures, SIGINT, cancelled shard), which say nothing about the test.
.filter((r) => r.status !== test.expectedStatus && (r.status === 'failed' || r.status === 'timedOut'))
.map(classify),
quarantine: parseQuarantine(test, now),
});
}
try {
fs.mkdirSync(path.dirname(this.outputFile), { recursive: true });
// Write-then-rename so a killed CI job never leaves half a JSON file
// that breaks the history aggregation later.
const tmp = this.outputFile + '.tmp';
fs.writeFileSync(tmp, JSON.stringify({ runStatus: result.status, entries }, null, 2));
fs.renameSync(tmp, this.outputFile);
} catch (err) {
// A reporter that throws hides the real test results. Log and move on.
console.error('[flaky-triage] could not write ' + this.outputFile + ':', err);
}
const flaky = entries.filter((e) => e.outcome === 'flaky');
const expired = entries.filter((e) => e.quarantine?.expired);
if (flaky.length > 0) {
console.log('[flaky-triage] ' + flaky.length + ' flaky: ' + flaky.map((e) => e.id).join(', '));
}
if (expired.length > 0) {
console.error(
'[flaky-triage] expired quarantine (fix or delete the test): ' +
expired.map((e) => e.id + ' [' + (e.quarantine?.ticket ?? 'no ticket') + ']').join(', '),
);
// Overriding the run status makes an expired quarantine block the build.
return { status: 'failed' };
}
// In the quarantine job, red tests are data, not a gate: report them in
// triage.json but exit 0, so only an expired quarantine can block a merge.
if (this.quarantineJob && result.status === 'failed') return { status: 'passed' };
}
}A few internals make this work. onTestEnd fires once per attempt, so the map accumulates every TestResult for a test before onEnd reads them. Using test.outcome() instead of the last result’s status means test.fail() tests that failed on purpose are reported as expected, and comparing each attempt to test.expectedStatus keeps their deliberate failures out of the signature list. The ANSI strip matters because Playwright colours its error messages, and an escape code in the middle of “Timeout” makes the regex miss. Finally, onEnd can return a status that overrides the run’s result. The reporter uses that in both directions: it fails the run when a quarantine has expired, and in the quarantine job it turns red tests into data instead of a blocking exit code.
We ran this reporter against a scratch suite before publishing: a test that throws on its first attempt was reported as flaky with one action-timeout signature, a test.fail() test came out expected with no signatures, and the quarantine job exited 0 with a failing quarantined test, then exited 1 once its expiry date was in the past.
Step 3: turn run history into a decision per test
One run tells you a test was flaky once. The matrix needs rates, so collect triage.json from every CI run on your main branch and aggregate. The script below reads a directory of those files, skips the ones it cannot trust, and prints a Markdown table you can paste into a ticket or a weekly quality review.
// scripts/triage-matrix.ts
// Usage: npx tsx scripts/triage-matrix.ts triage-history/
// Input: one triage.json per CI run on main (downloaded build artifacts).
import * as fs from 'node:fs';
import * as path from 'node:path';
import type { TriageEntry } from '../reporters/flaky-triage-reporter';
type Decision = 'insufficient-data' | 'fix-now' | 'quarantine' | 'retry-ok' | 'healthy';
// Illustrative starting points from our own pipelines, not industry
// standards. Tune them against your suite size and CI cost.
const MIN_RUNS = 10; // below this, one bad night looks like a 10%+ flake rate
const QUARANTINE_RATE = 0.05; // flaky in more than 1 run out of 20
const BROKEN_RATE = 0.3; // failing every retry this often is a regression, not noise
// Deterministic signatures: if these EVER flake, a retry is hiding a bug.
const ALWAYS_FIX = new Set(['strict-mode', 'detached-element', 'assertion']);
interface Stats {
runs: number;
flaky: number;
failed: number;
tags: Set<string>;
signatures: Map<string, number>;
}
function loadRuns(dir: string): TriageEntry[][] {
const runs: TriageEntry[][] = [];
for (const file of fs.readdirSync(dir).filter((f) => f.endsWith('.json'))) {
try {
const parsed = JSON.parse(fs.readFileSync(path.join(dir, file), 'utf8'));
if (!Array.isArray(parsed?.entries)) throw new Error('missing entries[]');
// A run that was interrupted or cancelled has no trustworthy outcomes.
if (parsed.runStatus === 'interrupted') {
console.warn('skip ' + file + ': interrupted run');
continue;
}
runs.push(parsed.entries);
} catch (err) {
// One truncated artifact must not abort the whole report.
console.warn('skip ' + file + ': ' + (err as Error).message);
}
}
return runs;
}
function decide(s: Stats): { decision: Decision; why: string } {
const flakeRate = s.flaky / s.runs;
const failRate = s.failed / s.runs;
const [topSig, topCount] = [...s.signatures].sort((a, b) => b[1] - a[1])[0] ?? ['none', 0];
if (s.runs < MIN_RUNS) return { decision: 'insufficient-data', why: s.runs + ' runs' };
if (failRate >= BROKEN_RATE) {
return { decision: 'fix-now', why: 'fails all retries in ' + Math.round(failRate * 100) + '% of runs' };
}
if (s.flaky + s.failed === 0) return { decision: 'healthy', why: 'no failed attempts' };
if (ALWAYS_FIX.has(topSig)) {
return { decision: 'fix-now', why: topSig + ' flakes are a race in the app or test, not infra noise' };
}
if (flakeRate >= QUARANTINE_RATE) {
// Critical-path tests are never quarantined: losing checkout coverage is
// worse than a red build. Escalate instead.
if (s.tags.has('@critical')) return { decision: 'fix-now', why: '@critical cannot be quarantined' };
return { decision: 'quarantine', why: Math.round(flakeRate * 100) + '% flaky, mostly ' + topSig };
}
const share = topCount / Math.max(1, s.flaky + s.failed);
return { decision: 'retry-ok', why: 'rare (' + s.flaky + '/' + s.runs + '), ' + topSig + ' ' + Math.round(share * 100) + '%' };
}
function main() {
const dir = process.argv[2];
if (!dir || !fs.existsSync(dir)) {
console.error('usage: triage-matrix <dir-with-triage-json-files>');
process.exit(2);
}
const runs = loadRuns(dir);
if (runs.length === 0) {
console.error('no readable triage files in ' + dir);
process.exit(2);
}
const byId = new Map<string, Stats>();
for (const entries of runs) {
for (const e of entries) {
if (e.outcome === 'skipped') continue; // a skip is not a pass
const s = byId.get(e.id) ?? { runs: 0, flaky: 0, failed: 0, tags: new Set(), signatures: new Map() };
s.runs++;
if (e.outcome === 'flaky') s.flaky++;
if (e.outcome === 'unexpected') s.failed++;
e.tags.forEach((t) => s.tags.add(t));
for (const sig of e.signatures) s.signatures.set(sig, (s.signatures.get(sig) ?? 0) + 1);
byId.set(e.id, s);
}
}
const rows = [...byId].map(([id, s]) => ({ id, ...decide(s) }));
const order: Decision[] = ['fix-now', 'quarantine', 'retry-ok', 'insufficient-data', 'healthy'];
rows.sort((a, b) => order.indexOf(a.decision) - order.indexOf(b.decision));
console.log('| Decision | Test | Why |\n|---|---|---|');
for (const r of rows.filter((r) => r.decision !== 'healthy')) {
console.log('| ' + r.decision + ' | ' + r.id.replace(/\|/g, '\\|') + ' | ' + r.why + ' |');
}
console.log('\n' + runs.length + ' runs, ' + byId.size + ' tests analysed.');
process.exitCode = rows.some((r) => r.decision === 'fix-now') ? 1 : 0;
}
main();Run against 12 synthetic reports (plus one deliberately truncated file), it produced this output. The data is illustrative, but the edge cases are real:
$ npx tsx scripts/triage-matrix.ts triage-history/
skip truncated.json: Unexpected end of JSON input
| Decision | Test | Why |
|---|---|---|
| fix-now | checkout.spec.ts › shows saved cards | strict-mode flakes are a race in the app or test, not infra noise |
| fix-now | checkout.spec.ts › updates VAT for EU address | fails all retries in 100% of runs |
| fix-now | cart.spec.ts › applies shipping | @critical cannot be quarantined |
| quarantine | checkout.spec.ts › applies a valid discount code | 8% flaky, mostly action-timeout |
12 runs, 5 tests analysed.Look at the last row. One flaky run out of 12 is 8%, which crosses the quarantine threshold. That is the small-sample problem in action: with 12 runs, a single bad night moves a test from “healthy” to “quarantine”. Raise MIN_RUNS or the window size before you automate anything on top of these decisions. The @critical row shows the other guard: a network flake on a critical path is escalated to a fix rather than hidden.
Wiring it into CI takes two changes: upload the triage file from every shard even when tests fail, and run the quarantine job separately so it never blocks merges for anything except an expired quarantine.
# .github/workflows/e2e.yml (excerpt)
jobs:
e2e:
runs-on: ubuntu-latest
strategy:
matrix:
shard: [1, 2, 3, 4]
steps:
- uses: actions/checkout@v4
- run: npm ci && npx playwright install --with-deps chromium
- run: npx playwright test --shard=${{ matrix.shard }}/4
# Upload even when tests fail, or you lose exactly the runs you need.
- if: ${{ !cancelled() }}
uses: actions/upload-artifact@v4
with:
name: triage-${{ github.run_id }}-shard-${{ matrix.shard }}
path: test-results/triage.json
quarantine:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npm ci && npx playwright install --with-deps chromium
# Exits 0 on red quarantined tests, 1 only on an expired quarantine.
- run: npx playwright test --repeat-each=3
env:
QUARANTINE_MODE: '1'Sharding needs no special handling in the script. Each shard writes its own file, every test lives in exactly one shard per run, so per-test counts stay correct.
When should you quarantine a flaky test instead of retrying it?
Answer capsule: Quarantine a non-critical test that flakes often enough to train people to ignore red builds, and only with an owner, a ticket, and an expiry.
A retry costs CI minutes and hides signal. A quarantine costs coverage: the behaviour that test protects is unguarded until someone fixes it. The trade is worth it when the flake is frequent enough that engineers start rerunning every red build by reflex, because at that point the test is not protecting anything anyway. It is never worth it for a test you would stop a release over.
Three rules keep quarantine from turning into a graveyard. Every quarantine names an owner and a ticket, so someone is accountable. Every quarantine has an expiry date that the reporter enforces, failing closed when the date is missing or unparseable. And quarantined tests keep running in their own job with --repeat-each, so when someone ships a fix you can see the flake rate drop to zero before the tag comes off, instead of guessing.
Step 4: fix the most common flake, the late listener
Async wait was the top root cause in the Luo et al. study, and in Playwright it most often looks like this: the test clicks a button, then starts waiting for the response the click triggered. On a slow runner, the listener is attached in time. On a fast one, the response has already arrived, and waitForResponse waits for a second one that never comes. For more patterns like this, see our deep dive into flaky test root causes.
// tests/checkout.spec.ts
import { test, expect } from '@playwright/test';
test.describe('checkout discounts', () => {
test.beforeEach(async ({ page }) => {
await page.goto('/cart?fixture=two-items');
await expect(page.getByTestId('cart-total')).toHaveText('€100.00');
});
// BEFORE (flaky): the listener is registered after the click. On a fast CI
// runner the response can arrive first, and waitForResponse then waits for
// a second response that never comes -> "Test timeout of 30000ms exceeded".
//
// await page.getByRole('button', { name: 'Apply code' }).click();
// await page.waitForResponse('**/api/cart/discount');
// const total = await page.getByTestId('cart-total').textContent();
// expect(total).toBe('€90.00'); // one-shot read, no auto-retry
test('applies a valid discount code', async ({ page }) => {
await page.getByLabel('Discount code').fill('AUTUMN10');
// Register the wait BEFORE the action that triggers the request. The
// predicate pins it to the POST carrying our code, so a debounced
// "validate as you type" GET to the same path cannot satisfy it.
const discountResponse = page.waitForResponse(
(res) =>
res.url().includes('/api/cart/discount') &&
res.request().method() === 'POST' &&
(res.request().postData() ?? '').includes('AUTUMN10'),
{ timeout: 10_000 },
);
await page.getByRole('button', { name: 'Apply code' }).click();
const res = await discountResponse;
// Fail with the server's reason instead of a generic UI timeout, so the
// triage reporter files this under "assertion", not "timeout".
if (!res.ok()) {
throw new Error('Discount API returned ' + res.status() + ': ' + (await res.text()));
}
// Web-first assertion: retries until the DOM settles or the expect timeout hits.
await expect(page.getByTestId('cart-total')).toHaveText('€90.00');
});
test(
'rejects an expired discount code',
{
tag: '@quarantine',
annotation: {
type: 'quarantine',
description: 'owner=@payments-qa; ticket=QA-412; expires=2026-10-15',
},
},
async ({ page }) => {
await page.getByLabel('Discount code').fill('SUMMER24');
await page.getByRole('button', { name: 'Apply code' }).click();
// Edge case: a stale success toast from a previous action can still be
// in the DOM. Scope the locator to the error role so strict mode stays happy.
await expect(page.getByRole('alert')).toHaveText(/expired/i);
await expect(page.getByTestId('cart-total')).toHaveText('€100.00');
},
);
});The fix works because page.waitForResponse returns a promise that subscribes to network events when it is called, not when it is awaited. Creating it before the click closes the window where the response can slip past. The predicate matters as much as the ordering: a glob like **/api/cart/discount also matches a debounced validation request fired while typing, which can satisfy the wait before the real POST happens.
| Flaky pattern | Deterministic pattern | Why the first one flakes |
|---|---|---|
await click(); await page.waitForResponse(url) | const p = page.waitForResponse(pred); await click(); await p | The response can arrive before the listener exists |
expect(await loc.textContent()).toBe('€90.00') | await expect(loc).toHaveText('€90.00') | A one-shot read captures the DOM mid-update; web-first assertions retry |
if (await loc.isVisible()) await loc.click() | await expect(loc).toBeVisible(); await loc.click() | isVisible() does not wait, so the branch depends on timing |
await page.waitForTimeout(2000) | await expect(page.getByRole('alert')).toHaveText(/saved/) | A fixed sleep is too long on fast runners and too short on slow ones |
Shared test@example.com user across tests | Per-test user seeded with testInfo.workerIndex or a fixture | Parallel workers mutate the same account and see each other’s data |
Troubleshooting: diagnosing a flake you cannot reproduce
A flake that never fails on your laptop is still diagnosable. Work through these in order; each one rules out a class of cause.
- Reproduce under load first. Run
npx playwright test tests/checkout.spec.ts --repeat-each=30 --retries=0 --workers=4. If it fails, rerun with--workers=1. A flake that disappears with one worker points at shared state (accounts, database rows, files), not timing. - Match CI’s CPU budget. Playwright defaults to half the machine’s logical CPU cores for workers. Your laptop and a two-core CI runner run very different amounts of work in parallel, and races that never fire locally fire constantly on a starved runner.
- Check order dependency. Run the failing test alone with
-g, then with its whole file, then with the shard it ran in. If it only fails after a specific neighbour, that neighbour leaks state. Remember that retries run in a fresh worker, which is exactly why order-dependent tests “heal” on retry. - Read the timeout correctly.
Test timeout of 30000ms exceededmeans the whole test ran out of budget, and the stuck step is in the trace.locator.click: Timeout 5000ms exceedednames the exact action, and the call log under it says why it could not act (not visible, covered by another element, not stable). - Open the failed attempt’s trace. With
retain-on-first-failure, the trace in the report is the failing run. Compare its network tab with a passing run: a response that lands before the action that should trigger it is the late-listener bug from Step 4. - Pin the clock and timezone. Tests that fail near midnight UTC or at month end are date bugs, not flakes. Set
timezoneIdinuseand control time withpage.clockinstead of trusting the runner’s clock.
For a longer walkthrough of trace-driven debugging, see our guide to debugging flaky tests.
Edge cases and gotchas
- Renames reset history. The reporter keys tests by title path, so renaming a test or moving it to another describe block starts a fresh history. A renamed flaky test shows up as
insufficient-data, not as fixed. - Skips are not passes. The history script ignores
skippedoutcomes. Counting them as runs would dilute the flake rate of a test that is conditionally skipped on some projects. - Interrupted runs are not evidence. When
maxFailuresstops a run or a CI job is cancelled, in-flight tests end asinterrupted. Both scripts drop them, otherwise a cancelled pipeline would show up as a burst of timeouts. - Expected failures stay out of the stats.
test.fail()tests fail on purpose. Comparing againstexpectedStatuskeeps them from looking permanently broken. - A reporter must never throw. An exception in
onEndcan mask the real test results. Catch, log, and let the run finish. - Upload artifacts on failure. If the upload step only runs on success, you collect history only for runs that had nothing interesting in them.
Putting it together
Retries are not the enemy. Unexamined retries are. Keep them on in CI, but make every flaky outcome produce a record: the attempt that failed, its signature, and the trace of the failure rather than the recovery. Aggregate those records across runs, and the retry-quarantine-fix question stops being a judgement call made at 6pm before a deploy. It becomes a table you review once a week: fix what the signatures say is a real race, quarantine what is noisy and non-critical with an owner and a deadline, and let the rare infrastructure blip retry in peace.
Start small. Add the reporter this week and collect two weeks of history before you act on any threshold. The first table it prints is usually enough to show which handful of tests account for most of your reruns.
Ready to strengthen your test automation?
Desplega.ai helps QA teams build robust test automation frameworks with reliable CI pipelines, flaky test triage, and quality gates that teams trust.
Get StartedFrequently Asked Questions
Should Playwright retries be enabled in CI at all?
Yes, as a detector. Keep retries at one or two in CI, record every flaky outcome with a reporter, and review the history weekly. Retries with no record just hide bugs.
How many runs do I need before calling a test flaky?
Collect at least 10 runs per test before acting, ideally more. With few runs a single bad night produces a huge rate, so treat a small sample as missing data, not as proof.
Is quarantining a flaky test the same as skipping it?
No. A quarantined test still runs in a separate non-blocking job, with an owner, a ticket, and an expiry. A skipped test produces no data, so nobody can tell when it is fixed.
Why does my flaky test only fail in CI and never locally?
CI runners usually have fewer cores and run more workers per core, so races that never fire locally fire there. Reproduce with --repeat-each and more workers, then compare.
Related Posts
When I Reject v0 Code: Pattern-Matching Rules for Safer UI Generation
A practical v0 review gate for safer generated React UI: AST checks, Playwright smoke tests, accessibility rules, and rejection signals.
Cody's Repository Indexing: Does Cognitive Offloading Create Knowledge Gaps in Large Codebases? | Desplega AI
A practical deep dive into Cody repository indexing, context retrieval, and how indie hackers avoid AI-created knowledge gaps.
Hot Module Replacement: Why Your Dev Server Restarts Are Killing Your Flow State | desplega.ai
Stop losing 2-3 hours daily to dev server restarts. Master HMR configuration in Vite and Next.js to maintain flow state, preserve component state, and boost coding velocity by 80%.