Skip to main content

Parallel Browser Execution Feature Plan

Status: Superseded

This is the original plan for parallel execution, kept for the reasoning behind the design. The feature shipped, and it shipped differently: the implementation uses a browser worker pool with a shared queue rather than the p-limit chunking this plan proposed.

Read BrowserWorkerPool for how parallel execution actually works, and Parallel Execution Performance Results for the measurements. Nothing below describes the current release.


Objective​

Implement parallel browser execution to run 4-8 browsers concurrently, reducing total test execution time by 2-4x.

Current state: Tests run sequentially (1 browser per project, projects run one after another) Target state: Tests run in parallel (N browsers, dynamic work distribution)


Background Research​

Current Implementation​

  • Location: packages/qs-design-tests/src/cli/executor.ts:237-330
  • Pattern: Sequential for loop over contexts (projects)
  • Browser lifecycle: Create browser → run all tests → close browser
  • Config field: playwright.workers already exists in schema (unused)

Proven Pattern: MCHWEB Implementation​

The old MCHWEB design tests used a custom BrowserWorkerPool:

  • Pre-allocates N browsers upfront
  • Simple polling-based availability check (getWorker() blocks until available)
  • Context reuse (closes old context, creates fresh one on return)
  • Critical insight: Serializes screenshot+diff operations to avoid disk I/O contention
  • Visual status logging: | 0 ✅ | 1 ❌ | 2 ✅ |

Performance results: 3-4x speedup with 4 workers (not perfect 4x due to overhead)

Playwright Best Practices​

  • Each worker should have its own browser instance (isolation)
  • Close contexts before browser for graceful cleanup
  • Browser reuse is efficient (contexts are cheap to create)

Implementation Approach​

Strategy: Option 3 (p-limit with Dynamic Load Balancing)​

Why this approach:

  1. Simplest to implement: ~30 lines of code with p-limit
  2. Dynamic work distribution: Workers pick up next test when done (efficient for uneven test durations)
  3. Lazy browser allocation: Only create browsers as needed
  4. Battle-tested library: p-limit is widely used and reliable
  5. Measure first, optimize later: If p-limit proves insufficient, upgrade to BrowserWorkerPool

Trade-offs:

  • ✅ Simpler than custom BrowserWorkerPool
  • ✅ Better load balancing than static Promise.all
  • ❌ Adds dependency (~5KB, acceptable)
  • ❌ No visual pool status logging (can add if needed)

Design Decisions​

1. Default Worker Count: 1 (Backwards Compatible)​

const workers = config.playwright?.workers || 1;

Rationale:

  • Safe default for CI environments with limited resources
  • Users opt-in to parallelism explicitly
  • Maintains current behavior unless configured otherwise

2. No Screenshot Serialization (Fully Parallel)​

Run screenshot capture and pixelmatch comparison in parallel without additional locking.

Rationale:

  • Measure performance first without optimization
  • Modern SSDs handle concurrent I/O well
  • If contention becomes an issue, add serialization later
  • Keep implementation simple initially

Note: MCHWEB serialized screenshot+diff operations (1 at a time) to avoid disk I/O contention. We'll skip this initially and measure. If performance degrades with high worker counts, we'll add a ScreenshotLock class.

3. p-limit Strategy (Not BrowserWorkerPool)​

Use p-limit for concurrency control, not custom pool class.

Rationale:

  • Faster to implement and test
  • Proven library with good API
  • Can upgrade to BrowserWorkerPool later if needed

Implementation Steps​

Step 1: Add p-limit Dependency​

cd packages/qs-design-tests
pnpm add p-limit

Package: p-limit@^5.0.0 (latest version)

Step 2: Modify Executor for Parallel Execution​

File: packages/qs-design-tests/src/cli/executor.ts

Current code (line 335-364):

export async function executeAllTests(
contexts: ProjectRunContext[],
config: DesignTestsConfig,
options: ExecutorOptions
): Promise<TestRunResults> {
// ...

// Execute each project sequentially
for (const context of contexts) {
const projectResults = await executeProjectTests(context, config, options, ...);
allResults.push(...projectResults);
}

// ...
}

New code (with p-limit):

import pLimit from 'p-limit';

export async function executeAllTests(
contexts: ProjectRunContext[],
config: DesignTestsConfig,
options: ExecutorOptions
): Promise<TestRunResults> {
const startTime = Date.now();
const allResults: TestResult[] = [];

const totalTests = contexts.reduce((sum, ctx) => sum + ctx.testSpecs.length, 0);
const totalStories = new Set(contexts.flatMap((ctx) => ctx.testSpecs.map((s) => s.story.id))).size;
const totalViewports = new Set(contexts.flatMap((ctx) => ctx.testSpecs.map((s) => s.viewport.name))).size;

logTestStartSummary(totalTests, totalStories, totalViewports, contexts.length);

// Get worker count from config (default: 1 for backwards compatibility)
const workers = config.playwright?.workers || 1;

if (workers === 1) {
// Sequential execution (current behavior)
console.log("🔄 Running tests sequentially (1 worker)");

for (const context of contexts) {
const projectResults = await executeProjectTests(
context,
config,
options,
(current, total, result, retry) => {
logTestProgress(current, total, result.viewport, result.story, retry);
}
);
allResults.push(...projectResults);
}
} else {
// Parallel execution with p-limit
console.log(`⚡ Running tests in parallel (${workers} workers)`);

const limit = pLimit(workers);

const parallelResults = await Promise.all(
contexts.map((context) =>
limit(async () => {
console.log(`🚀 Worker starting project: ${context.projectKey}`);
return executeProjectTests(
context,
config,
options,
(current, total, result, retry) => {
logTestProgress(current, total, result.viewport, result.story, retry);
}
);
})
)
);

allResults.push(...parallelResults.flat());
}

console.log(); // Final newline

const duration = Date.now() - startTime;
return aggregateResults(allResults, duration);
}

Key changes:

  1. Check config.playwright?.workers (defaults to 1)
  2. Branch on workers === 1 (sequential) vs > 1 (parallel)
  3. Use pLimit(workers) to create concurrency limiter
  4. Wrap each executeProjectTests() call in limit()
  5. Use Promise.all() to wait for all tests to complete
  6. Flatten results with .flat() (array of arrays → flat array)

Step 3: Add Logging for Parallel Mode​

Current logging:

🧪 Starting tests: 14 total, 7 stories, 2 viewports, 1 project

Enhanced logging for parallel:

🧪 Starting tests: 14 total, 7 stories, 2 viewports, 1 project
⚡ Running tests in parallel (4 workers)
🚀 Worker starting project: bfh
🚀 Worker starting project: hkb
...

Step 4: Wire Up Configuration​

Already works! The config schema has playwright.workers field:

// packages/qs-design-tests/src/config/schema.ts:81-87
playwright: z.object({
workers: z.number().positive().optional(),
timeout: z.number().positive().optional(),
retries: z.number().nonnegative().optional(),
headless: z.boolean().optional(),
ignoreHTTPSErrors: z.boolean().optional(),
proxy: z.object({ server: z.string() }).optional(),
}).optional(),

Configuration methods:

# 1. CLI flag
pnpm exec design-tests run --config .designTests.js --all --set playwright.workers=4

# 2. Environment variable
DESIGN_TESTS__PLAYWRIGHT__WORKERS=4 pnpm test:bfh:ci

# 3. Config file (.designTests.js)
export default {
playwright: {
workers: 4, // 4 browsers in parallel
headless: true,
},
// ...
}

Step 5: Handle Browser Lifecycle​

Current implementation in executeProjectTests() (executor.ts:237-330):

  • Creates 1 browser per project
  • Closes browser after all tests complete

For parallel execution:

  • Each call to executeProjectTests() creates its own browser
  • If 4 workers and 4 projects: 4 browsers total (1 per worker)
  • If 4 workers and 10 projects: max 4 browsers at a time (p-limit ensures this)

No changes needed - existing browser lifecycle already works correctly for parallel execution!


Testing Plan​

Phase 1: Correctness Testing​

Goal: Verify parallel execution produces identical results to sequential.

Test cases:

  1. Baseline (sequential): workers=1
  2. Small parallel: workers=2
  3. Medium parallel: workers=4
  4. Large parallel: workers=8

Verification:

  • Compare actual screenshots pixel-by-pixel
  • Verify no race conditions (file conflicts, etc.)
  • Check all tests complete successfully
  • Ensure no flakiness (run 3 times each)

Command:

cd packages/qs-design-tests

# Baseline
DESIGN_TESTS__PLAYWRIGHT__WORKERS=1 time node dist/cli.js run --config .designTests.js --all

# Parallel (4 workers)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=4 time node dist/cli.js run --config .designTests.js --all

# Parallel (8 workers)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=8 time node dist/cli.js run --config .designTests.js --all

Phase 2: Performance Benchmarking​

Test environment:

  • MacBook Pro (local)
  • BFH config: 7 pages × 2 viewports = 14 tests
  • Headless mode
  • Localhost Storybook (fast network)

Metrics to measure:

  • Total duration (wall clock time)
  • Per-test average duration
  • Browser launch overhead
  • CPU/memory usage (optional)

Expected results:

Workers | Duration | Speedup | Notes
--------|----------|---------|------
1 | ~120s | 1.0x | Baseline (sequential)
2 | ~70s | 1.7x | Good scaling
4 | ~40s | 3.0x | Near-ideal (some overhead)
8 | ~30s | 4.0x | Diminishing returns

Success criteria:

  • 4 workers: ≥2.5x speedup (acceptable)
  • 8 workers: ≥3.0x speedup (good)
  • No visual regressions
  • No flakiness

Phase 3: CI Environment Testing​

Test in Jenkins:

# Modify Jenkinsfile environment block
environment {
DESIGN_TESTS__PLAYWRIGHT__WORKERS = '4'
}

Verify:

  • CI performance improvement
  • K8s resource limits not exceeded
  • No OOM errors
  • Artifacts correctly generated

Rollout Plan​

Stage 1: Local Development Testing (1 day)​

  1. Implement p-limit solution
  2. Test locally with BFH config (14 tests)
  3. Verify correctness (sequential vs parallel comparison)
  4. Measure performance (1 vs 4 vs 8 workers)

Stage 2: CI Testing (1 day)​

  1. Enable 4 workers in Jenkins dev pipeline
  2. Monitor build duration
  3. Check for resource issues (CPU, memory)
  4. Verify artifacts and reports

Stage 3: Production Rollout (phased)​

  1. Enable 4 workers in production pipeline
  2. Monitor for 1 week
  3. Gradually increase to 8 workers if stable
  4. Document best practices

Stage 4: Decide on BrowserWorkerPool​

Decision point: After measuring p-limit performance

Upgrade to BrowserWorkerPool if:

  • p-limit doesn't scale beyond 4 workers
  • Need better visibility (pool status logging)
  • Need more control (custom retry logic, etc.)

Keep p-limit if:

  • 3-4x speedup achieved with 4 workers
  • No stability issues
  • Simpler is better

Success Criteria​

Must Have (P0)​

  • ✅ 2.5x speedup with 4 workers
  • ✅ No visual regressions (pixel-perfect match)
  • ✅ No race conditions or file conflicts
  • ✅ Backwards compatible (workers=1 is default)
  • ✅ Works in CI (Jenkins K8s)

Nice to Have (P1)​

  • Better logging (show which worker is doing what)
  • Pool status visualization (like MCHWEB)
  • Per-worker timing statistics
  • Auto-detect optimal worker count (CPU cores)

Future Enhancements (P2)​

  • BrowserWorkerPool implementation (if p-limit insufficient)
  • Screenshot serialization (if I/O contention observed)
  • Smart work distribution (prioritize slow tests)

Risk Assessment​

High Risk ⚠️​

Race conditions in filesystem operations

  • Multiple browsers writing to same directories
  • Mitigation: Test thoroughly, add file locks if needed

OOM in CI (Jenkins K8s pod)

  • 8 browsers × 200MB RAM = 1.6GB (pod has 6GB limit)
  • Mitigation: Start with 4 workers, monitor memory usage

Medium Risk ⚠️​

Diminishing returns beyond 4 workers

  • I/O bound (disk writes, pixelmatch CPU)
  • Mitigation: Benchmark and document optimal worker count

Flakiness increase

  • More parallelism = more complexity
  • Mitigation: Run multiple times, verify stability

Low Risk ✅​

p-limit dependency

  • Well-maintained, widely used (30M+ downloads/week)
  • Mitigation: Pin version, test thoroughly

Implementation Checklist​

Code Changes​

  • Add p-limit to package.json dependencies
  • Modify executor.ts:executeAllTests() for parallel execution
  • Add logging for parallel mode (worker count, project assignments)
  • Update type definitions if needed

Testing​

  • Local correctness testing (sequential vs parallel)
  • Performance benchmarking (1, 4, 8 workers)
  • CI integration testing (Jenkins)
  • Stress test with full MCHWEB config (100+ tests)

Documentation​

  • Update README.md with worker configuration examples
  • Update CLAUDE.md with architecture details
  • Create parallel-execution-results.md benchmark report (already exists)
  • Add JSDoc comments to new code

Deployment​

  • Merge to develop branch
  • Deploy to Jenkins dev pipeline
  • Monitor for 3 days
  • Enable in production pipeline

Configuration Examples​

Local Development​

# 4 workers (recommended)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=4 pnpm test:bfh

# 8 workers (max)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=8 pnpm test:bfh

# Sequential (debugging)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=1 pnpm test:bfh

Jenkins CI​

// In Jenkinsfile environment block
environment {
DESIGN_TESTS__PLAYWRIGHT__WORKERS = '4' // 4 parallel browsers
}

Config File​

// .designTests.js
export default {
playwright: {
workers: 4, // Run 4 browsers in parallel
headless: true,
timeout: 120000,
},
// ...
};

Follow-Up Work​

If p-limit Works Well ✅​

  1. Document best practices (optimal worker count per environment)
  2. Add to what-design-tests-are.md for stakeholders
  3. Roll out to other projects (MCHWEB, EWZ)
  4. Consider making 4 workers the default (after monitoring)

If p-limit Insufficient ⚠️​

  1. Implement BrowserWorkerPool (Option 1)
  2. Add visual pool status logging
  3. Add screenshot serialization (if I/O contention observed)
  4. Add custom retry logic per worker
  5. Document lessons learned

References​

Code Locations​

  • Current executor: packages/qs-design-tests/src/cli/executor.ts:335-371
  • Config schema: packages/qs-design-tests/src/config/schema.ts:81-87
  • MCHWEB BrowserWorkerPool: /Users/ma/CODE/MCHWEB/mchweb-design-tests/tests/BrowserWorkerPool.ts

External Resources​


Last Updated: 2025-10-22 Author: Max Albrecht (with Claude Code) Next Review: After Stage 1 completion (local testing)