Parallel Browser Execution Feature Plan
Status: Superseded
This is the original plan for parallel execution, kept for the reasoning behind the design. The feature shipped, and it shipped differently: the implementation uses a browser worker pool with a shared queue rather than the p-limit chunking this plan proposed.
Read BrowserWorkerPool for how parallel execution actually works, and Parallel Execution Performance Results for the measurements. Nothing below describes the current release.
Objective
Implement parallel browser execution to run 4-8 browsers concurrently, reducing total test execution time by 2-4x.
Current state: Tests run sequentially (1 browser per project, projects run one after another) Target state: Tests run in parallel (N browsers, dynamic work distribution)
Background Research
Current Implementation
- Location:
packages/qs-design-tests/src/cli/executor.ts:237-330 - Pattern: Sequential
forloop overcontexts(projects) - Browser lifecycle: Create browser → run all tests → close browser
- Config field:
playwright.workersalready exists in schema (unused)
Proven Pattern: MCHWEB Implementation
The old MCHWEB design tests used a custom BrowserWorkerPool:
- Pre-allocates N browsers upfront
- Simple polling-based availability check (
getWorker()blocks until available) - Context reuse (closes old context, creates fresh one on return)
- Critical insight: Serializes screenshot+diff operations to avoid disk I/O contention
- Visual status logging:
| 0 ✅ | 1 ❌ | 2 ✅ |
Performance results: 3-4x speedup with 4 workers (not perfect 4x due to overhead)
Playwright Best Practices
- Each worker should have its own browser instance (isolation)
- Close contexts before browser for graceful cleanup
- Browser reuse is efficient (contexts are cheap to create)
Implementation Approach
Strategy: Option 3 (p-limit with Dynamic Load Balancing)
Why this approach:
- Simplest to implement: ~30 lines of code with p-limit
- Dynamic work distribution: Workers pick up next test when done (efficient for uneven test durations)
- Lazy browser allocation: Only create browsers as needed
- Battle-tested library: p-limit is widely used and reliable
- Measure first, optimize later: If p-limit proves insufficient, upgrade to BrowserWorkerPool
Trade-offs:
- ✅ Simpler than custom BrowserWorkerPool
- ✅ Better load balancing than static Promise.all
- ❌ Adds dependency (~5KB, acceptable)
- ❌ No visual pool status logging (can add if needed)
Design Decisions
1. Default Worker Count: 1 (Backwards Compatible)
const workers = config.playwright?.workers || 1;
Rationale:
- Safe default for CI environments with limited resources
- Users opt-in to parallelism explicitly
- Maintains current behavior unless configured otherwise
2. No Screenshot Serialization (Fully Parallel)
Run screenshot capture and pixelmatch comparison in parallel without additional locking.
Rationale:
- Measure performance first without optimization
- Modern SSDs handle concurrent I/O well
- If contention becomes an issue, add serialization later
- Keep implementation simple initially
Note: MCHWEB serialized screenshot+diff operations (1 at a time) to avoid disk I/O contention. We'll skip this initially and measure. If performance degrades with high worker counts, we'll add a ScreenshotLock class.
3. p-limit Strategy (Not BrowserWorkerPool)
Use p-limit for concurrency control, not custom pool class.
Rationale:
- Faster to implement and test
- Proven library with good API
- Can upgrade to BrowserWorkerPool later if needed
Implementation Steps
Step 1: Add p-limit Dependency
cd packages/qs-design-tests
pnpm add p-limit
Package: p-limit@^5.0.0 (latest version)
Step 2: Modify Executor for Parallel Execution
File: packages/qs-design-tests/src/cli/executor.ts
Current code (line 335-364):
export async function executeAllTests(
contexts: ProjectRunContext[],
config: DesignTestsConfig,
options: ExecutorOptions
): Promise<TestRunResults> {
// ...
// Execute each project sequentially
for (const context of contexts) {
const projectResults = await executeProjectTests(context, config, options, ...);
allResults.push(...projectResults);
}
// ...
}
New code (with p-limit):
import pLimit from 'p-limit';
export async function executeAllTests(
contexts: ProjectRunContext[],
config: DesignTestsConfig,
options: ExecutorOptions
): Promise<TestRunResults> {
const startTime = Date.now();
const allResults: TestResult[] = [];
const totalTests = contexts.reduce((sum, ctx) => sum + ctx.testSpecs.length, 0);
const totalStories = new Set(contexts.flatMap((ctx) => ctx.testSpecs.map((s) => s.story.id))).size;
const totalViewports = new Set(contexts.flatMap((ctx) => ctx.testSpecs.map((s) => s.viewport.name))).size;
logTestStartSummary(totalTests, totalStories, totalViewports, contexts.length);
// Get worker count from config (default: 1 for backwards compatibility)
const workers = config.playwright?.workers || 1;
if (workers === 1) {
// Sequential execution (current behavior)
console.log("🔄 Running tests sequentially (1 worker)");
for (const context of contexts) {
const projectResults = await executeProjectTests(
context,
config,
options,
(current, total, result, retry) => {
logTestProgress(current, total, result.viewport, result.story, retry);
}
);
allResults.push(...projectResults);
}
} else {
// Parallel execution with p-limit
console.log(`⚡ Running tests in parallel (${workers} workers)`);
const limit = pLimit(workers);
const parallelResults = await Promise.all(
contexts.map((context) =>
limit(async () => {
console.log(`🚀 Worker starting project: ${context.projectKey}`);
return executeProjectTests(
context,
config,
options,
(current, total, result, retry) => {
logTestProgress(current, total, result.viewport, result.story, retry);
}
);
})
)
);
allResults.push(...parallelResults.flat());
}
console.log(); // Final newline
const duration = Date.now() - startTime;
return aggregateResults(allResults, duration);
}
Key changes:
- Check
config.playwright?.workers(defaults to 1) - Branch on workers === 1 (sequential) vs > 1 (parallel)
- Use
pLimit(workers)to create concurrency limiter - Wrap each
executeProjectTests()call inlimit() - Use
Promise.all()to wait for all tests to complete - Flatten results with
.flat()(array of arrays → flat array)
Step 3: Add Logging for Parallel Mode
Current logging:
🧪 Starting tests: 14 total, 7 stories, 2 viewports, 1 project
Enhanced logging for parallel:
🧪 Starting tests: 14 total, 7 stories, 2 viewports, 1 project
⚡ Running tests in parallel (4 workers)
🚀 Worker starting project: bfh
🚀 Worker starting project: hkb
...
Step 4: Wire Up Configuration
Already works! The config schema has playwright.workers field:
// packages/qs-design-tests/src/config/schema.ts:81-87
playwright: z.object({
workers: z.number().positive().optional(),
timeout: z.number().positive().optional(),
retries: z.number().nonnegative().optional(),
headless: z.boolean().optional(),
ignoreHTTPSErrors: z.boolean().optional(),
proxy: z.object({ server: z.string() }).optional(),
}).optional(),
Configuration methods:
# 1. CLI flag
pnpm exec design-tests run --config .designTests.js --all --set playwright.workers=4
# 2. Environment variable
DESIGN_TESTS__PLAYWRIGHT__WORKERS=4 pnpm test:bfh:ci
# 3. Config file (.designTests.js)
export default {
playwright: {
workers: 4, // 4 browsers in parallel
headless: true,
},
// ...
}
Step 5: Handle Browser Lifecycle
Current implementation in executeProjectTests() (executor.ts:237-330):
- Creates 1 browser per project
- Closes browser after all tests complete
For parallel execution:
- Each call to
executeProjectTests()creates its own browser - If 4 workers and 4 projects: 4 browsers total (1 per worker)
- If 4 workers and 10 projects: max 4 browsers at a time (p-limit ensures this)
No changes needed - existing browser lifecycle already works correctly for parallel execution!
Testing Plan
Phase 1: Correctness Testing
Goal: Verify parallel execution produces identical results to sequential.
Test cases:
- Baseline (sequential):
workers=1 - Small parallel:
workers=2 - Medium parallel:
workers=4 - Large parallel:
workers=8
Verification:
- Compare actual screenshots pixel-by-pixel
- Verify no race conditions (file conflicts, etc.)
- Check all tests complete successfully
- Ensure no flakiness (run 3 times each)
Command:
cd packages/qs-design-tests
# Baseline
DESIGN_TESTS__PLAYWRIGHT__WORKERS=1 time node dist/cli.js run --config .designTests.js --all
# Parallel (4 workers)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=4 time node dist/cli.js run --config .designTests.js --all
# Parallel (8 workers)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=8 time node dist/cli.js run --config .designTests.js --all
Phase 2: Performance Benchmarking
Test environment:
- MacBook Pro (local)
- BFH config: 7 pages × 2 viewports = 14 tests
- Headless mode
- Localhost Storybook (fast network)
Metrics to measure:
- Total duration (wall clock time)
- Per-test average duration
- Browser launch overhead
- CPU/memory usage (optional)
Expected results:
Workers | Duration | Speedup | Notes
--------|----------|---------|------
1 | ~120s | 1.0x | Baseline (sequential)
2 | ~70s | 1.7x | Good scaling
4 | ~40s | 3.0x | Near-ideal (some overhead)
8 | ~30s | 4.0x | Diminishing returns
Success criteria:
- 4 workers: ≥2.5x speedup (acceptable)
- 8 workers: ≥3.0x speedup (good)
- No visual regressions
- No flakiness
Phase 3: CI Environment Testing
Test in Jenkins:
# Modify Jenkinsfile environment block
environment {
DESIGN_TESTS__PLAYWRIGHT__WORKERS = '4'
}
Verify:
- CI performance improvement
- K8s resource limits not exceeded
- No OOM errors
- Artifacts correctly generated
Rollout Plan
Stage 1: Local Development Testing (1 day)
- Implement p-limit solution
- Test locally with BFH config (14 tests)
- Verify correctness (sequential vs parallel comparison)
- Measure performance (1 vs 4 vs 8 workers)
Stage 2: CI Testing (1 day)
- Enable 4 workers in Jenkins dev pipeline
- Monitor build duration
- Check for resource issues (CPU, memory)
- Verify artifacts and reports
Stage 3: Production Rollout (phased)
- Enable 4 workers in production pipeline
- Monitor for 1 week
- Gradually increase to 8 workers if stable
- Document best practices
Stage 4: Decide on BrowserWorkerPool
Decision point: After measuring p-limit performance
Upgrade to BrowserWorkerPool if:
- p-limit doesn't scale beyond 4 workers
- Need better visibility (pool status logging)
- Need more control (custom retry logic, etc.)
Keep p-limit if:
- 3-4x speedup achieved with 4 workers
- No stability issues
- Simpler is better
Success Criteria
Must Have (P0)
- ✅ 2.5x speedup with 4 workers
- ✅ No visual regressions (pixel-perfect match)
- ✅ No race conditions or file conflicts
- ✅ Backwards compatible (workers=1 is default)
- ✅ Works in CI (Jenkins K8s)
Nice to Have (P1)
- Better logging (show which worker is doing what)
- Pool status visualization (like MCHWEB)
- Per-worker timing statistics
- Auto-detect optimal worker count (CPU cores)
Future Enhancements (P2)
- BrowserWorkerPool implementation (if p-limit insufficient)
- Screenshot serialization (if I/O contention observed)
- Smart work distribution (prioritize slow tests)
Risk Assessment
High Risk ⚠️
Race conditions in filesystem operations
- Multiple browsers writing to same directories
- Mitigation: Test thoroughly, add file locks if needed
OOM in CI (Jenkins K8s pod)
- 8 browsers × 200MB RAM = 1.6GB (pod has 6GB limit)
- Mitigation: Start with 4 workers, monitor memory usage
Medium Risk ⚠️
Diminishing returns beyond 4 workers
- I/O bound (disk writes, pixelmatch CPU)
- Mitigation: Benchmark and document optimal worker count
Flakiness increase
- More parallelism = more complexity
- Mitigation: Run multiple times, verify stability
Low Risk ✅
p-limit dependency
- Well-maintained, widely used (30M+ downloads/week)
- Mitigation: Pin version, test thoroughly
Implementation Checklist
Code Changes
- Add
p-limitto package.json dependencies - Modify
executor.ts:executeAllTests()for parallel execution - Add logging for parallel mode (worker count, project assignments)
- Update type definitions if needed
Testing
- Local correctness testing (sequential vs parallel)
- Performance benchmarking (1, 4, 8 workers)
- CI integration testing (Jenkins)
- Stress test with full MCHWEB config (100+ tests)
Documentation
- Update
README.mdwith worker configuration examples - Update
CLAUDE.mdwith architecture details - Create
parallel-execution-results.mdbenchmark report (already exists) - Add JSDoc comments to new code
Deployment
- Merge to develop branch
- Deploy to Jenkins dev pipeline
- Monitor for 3 days
- Enable in production pipeline
Configuration Examples
Local Development
# 4 workers (recommended)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=4 pnpm test:bfh
# 8 workers (max)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=8 pnpm test:bfh
# Sequential (debugging)
DESIGN_TESTS__PLAYWRIGHT__WORKERS=1 pnpm test:bfh
Jenkins CI
// In Jenkinsfile environment block
environment {
DESIGN_TESTS__PLAYWRIGHT__WORKERS = '4' // 4 parallel browsers
}
Config File
// .designTests.js
export default {
playwright: {
workers: 4, // Run 4 browsers in parallel
headless: true,
timeout: 120000,
},
// ...
};
Follow-Up Work
If p-limit Works Well ✅
- Document best practices (optimal worker count per environment)
- Add to what-design-tests-are.md for stakeholders
- Roll out to other projects (MCHWEB, EWZ)
- Consider making 4 workers the default (after monitoring)
If p-limit Insufficient ⚠️
- Implement BrowserWorkerPool (Option 1)
- Add visual pool status logging
- Add screenshot serialization (if I/O contention observed)
- Add custom retry logic per worker
- Document lessons learned
References
Code Locations
- Current executor:
packages/qs-design-tests/src/cli/executor.ts:335-371 - Config schema:
packages/qs-design-tests/src/config/schema.ts:81-87 - MCHWEB BrowserWorkerPool:
/Users/ma/CODE/MCHWEB/mchweb-design-tests/tests/BrowserWorkerPool.ts
External Resources
- p-limit docs: https://www.npmjs.com/package/p-limit
- Playwright parallelism: https://playwright.dev/docs/test-parallel
- Playwright Browser API: https://playwright.dev/docs/api/class-browser
Last Updated: 2025-10-22 Author: Max Albrecht (with Claude Code) Next Review: After Stage 1 completion (local testing)