BreakYourApp.

BREAKYOURAPP · TESTING NOTES

BreakYourApp QA Benchmark: Five Seeded Defects, Fourteen Runs

Published · About BreakYourApp

On October 3, 2026, our production browser and assertion components ran against a disposable SaaS fixture with seven scenarios, each repeated twice. All five deliberately introduced defects produced the expected failed check. This validates the configured checks in this fixture; it is not a customer accuracy claim.

The fixture and the actual test path

One loopback application provides login, a project-creation workflow, a private UI resource and a JSON API. It uses synthetic accounts A and B and disposable data. Each run creates fresh browser sessions; the fixture is not published as a vulnerable public website.

The test uses the production BrowserAgent, executeJourney and runAccessChecks implementations. The configured journey signs in, creates one project, checks the declared creation request, counts the expected result and checks the result after refresh. Account-isolation checks establish A’s baseline, verify B’s authenticated marker and request the private resource as B and anonymously.

Expected versus observed outcomes

The healthy control passed both the configured journey and private API comparison. Five defect scenarios failed the relevant check in both repetitions. In the failed-login scenario, the journey for A passed while B’s API comparison was correctly blocked.

Expected versus observed outcomes: test conditions and evidence
Test conditionExpected observationEvidence or interpretation
Healthy applicationJourney passed; API isolation passed2 of 2 expected results
Lost state after refreshJourney failed2 of 2 detected
Double-submit duplicateJourney failed2 of 2 detected
HTTP 500 with UI successJourney failed2 of 2 detected
Private JSON exposed despite denied UIAPI check failed2 of 2 detected
403 containing private JSONAPI check failed2 of 2 detected
Failed B loginB API check blocked2 of 2 expected results

Results and interpretation

The suite completed 14 runs covering seven scenarios and five seeded defects. It reported five of five detected defects, zero false failures and zero false passes across the specific expected journey and secondary-API verdicts. The suite also checked the anonymous API verdict against its expected result.

These are small, controlled measurements. They do not estimate the percentage of unknown bugs the product will find in a customer app. The test supplies labels, URLs, expected values and policy context. An AI-generated test plan was not responsible for discovering the fixture’s scenarios.

Reproduce and inspect the evidence

The repository contains scripts/saas-simulation.ts. With the repository’s dependencies and Playwright Chromium installed, run npm run benchmark:saas. The program exits unsuccessfully if an expected verdict differs.

The run writes simulation-results/saas-report.json and saas-report.md with expected and actual verdicts, journey checks and access-check reasons. GitHub Actions saves both as the saas-simulation-results artifact. The benchmark also verifies that the synthetic private marker is absent from persisted access-check results.

What has not been validated here

The simulation does not exercise autonomous AI discovery, the scan queue, deployed customer infrastructure, email delivery or PDF rendering. Refresh persistence is a UI observation in this fixture, not proof of durability across sessions or database failures.

Further validation should run the complete product on an authorized customer staging environment, compare findings with human QA, and measure missed defects, useful findings, setup effort and the complete report experience. Repeating this benchmark in CI helps detect regressions in the checks it actually covers.

Inspect the October 3 simulation run and downloadable results

Related testing guides