SauceDemo 55% 6/10 scenarios · Stable 4 completion blockers
Sauce Shopify 42.9% 4/8 scenarios · Not stable 6 completion blockers
Automation Exercise 33.3% 5/11 scenarios · Stable 6 completion blockers

Portfolio: 43.9% weighted coverage · 15/29 scenarios · 2/3 projects stable

Live CI telemetry

30-Day QA Health

Waiting for the first published agent run.

-- Regression tests passing
-- Passing agent candidates
-- Successful repaired runs
-- Runs with timeouts
Application Regression reliability Candidate success Timeouts Latest result
Waiting for published run data.

Recent Activity

Latest published runs

No live runs published yet

The scheduled agent workflow will populate this section automatically.

Current coverage

Web Applications

Permanent regression suites and agent-generated candidates are grouped by application.

How it works

Agent Pipeline

  1. ExplorePlaywright opens the target page and the model chooses safe next actions.
  2. GenerateThe model turns the exploration trace into Playwright tests.
  3. ExecuteThe generated spec runs in Chromium with traces and screenshots on failure.
  4. RepairFailures are sent back to the model for selector and assertion repair.
  5. ReportRuns produce dashboard data and Playwright reports for review.

Next steps

Portfolio Roadmap

Live Report Publishing

Publish the latest dashboard and report artifacts to a stable URL.

Candidate Promotion

Review validated agent candidates and promote useful coverage into the permanent regression suite.

Specialized Agents

Split exploration, repair, accessibility, and reporting into focused roles when complexity grows.