BREAKYOURAPP · TESTING NOTES
BreakYourApp QA Benchmark: Five Seeded Defects, Fourteen Runs
Published · About BreakYourApp
On October 3, 2026, our production browser and assertion components ran against a disposable SaaS fixture with seven scenarios, each repeated twice. All five deliberately introduced defects produced the expected failed check. This validates the configured checks in this fixture; it is not a customer accuracy claim.
The fixture and the actual test path
One loopback application provides login, a project-creation workflow, a private UI resource and a JSON API. It uses synthetic accounts A and B and disposable data. Each run creates fresh browser sessions; the fixture is not published as a vulnerable public website.
The test uses the production BrowserAgent, executeJourney and runAccessChecks implementations. The configured journey signs in, creates one project, checks the declared creation request, counts the expected result and checks the result after refresh. Account-isolation checks establish A’s baseline, verify B’s authenticated marker and request the private resource as B and anonymously.
Expected versus observed outcomes
The healthy control passed both the configured journey and private API comparison. Five defect scenarios failed the relevant check in both repetitions. In the failed-login scenario, the journey for A passed while B’s API comparison was correctly blocked.
| Test condition | Expected observation | Evidence or interpretation |
|---|---|---|
| Healthy application | Journey passed; API isolation passed | 2 of 2 expected results |
| Lost state after refresh | Journey failed | 2 of 2 detected |
| Double-submit duplicate | Journey failed | 2 of 2 detected |
| HTTP 500 with UI success | Journey failed | 2 of 2 detected |
| Private JSON exposed despite denied UI | API check failed | 2 of 2 detected |
| 403 containing private JSON | API check failed | 2 of 2 detected |
| Failed B login | B API check blocked | 2 of 2 expected results |
Results and interpretation
The suite completed 14 runs covering seven scenarios and five seeded defects. It reported five of five detected defects, zero false failures and zero false passes across the specific expected journey and secondary-API verdicts. The suite also checked the anonymous API verdict against its expected result.
These are small, controlled measurements. They do not estimate the percentage of unknown bugs the product will find in a customer app. The test supplies labels, URLs, expected values and policy context. An AI-generated test plan was not responsible for discovering the fixture’s scenarios.
Reproduce and inspect the evidence
The repository contains scripts/saas-simulation.ts. With the repository’s dependencies and Playwright Chromium installed, run npm run benchmark:saas. The program exits unsuccessfully if an expected verdict differs.
The run writes simulation-results/saas-report.json and saas-report.md with expected and actual verdicts, journey checks and access-check reasons. GitHub Actions saves both as the saas-simulation-results artifact. The benchmark also verifies that the synthetic private marker is absent from persisted access-check results.
What has not been validated here
The simulation does not exercise autonomous AI discovery, the scan queue, deployed customer infrastructure, email delivery or PDF rendering. Refresh persistence is a UI observation in this fixture, not proof of durability across sessions or database failures.
Further validation should run the complete product on an authorized customer staging environment, compare findings with human QA, and measure missed defects, useful findings, setup effort and the complete report experience. Repeating this benchmark in CI helps detect regressions in the checks it actually covers.
Inspect the October 3 simulation run and downloadable results