Measured, not claimed
The month our engine finally beat the chat model
We built a test-generation engine, measured it honestly against a chat model, and lost. Here is the number we published while losing, what we changed, and the number now.
10
bugs seeded into a working store, one broken rule each
3 → 5
bugs our engine caught, before and after the repair loop
6 → 3
false alarms on a store with nothing wrong with it
4
caught by a chat model handed every page of the site
Losing in public
The first run was not close in the direction we wanted. Our engine — which explores the app in a real browser before writing anything — caught 3 of 10 seeded bugs. A chat model handed the text of every page caught 4. Worse, our suite raised 10 false alarms on the working store, and all three of our catches came from the deterministic layer that uses no AI at all.
The diagnosis was specific and unflattering. Our explorer followed links but never clicked anything, so it never saw a cart with something in it. The model, writing against that half-blind map, assumed “Add to cart” navigated to the cart — and every cart scenario failed on a store where nothing was wrong.
How the benchmark works
Build a store whose rules are written down
A small shop — search, cart, promo codes, checkout — where every business rule is stated on its own pages: one promo code per order, FLAT5 needs a $30 subtotal, up to 10 of any item, ZIP must be five digits.
Break exactly one rule at a time
Ten copies of the store, each with a single rule quietly broken. This is mutation testing: a test suite is only worth what it notices, so we measure suites by how many of the ten they kill.
Let a chat model compete properly
The same model, the same feature descriptions — and, for the strongest baseline, the full visible text of every page of the store. No trick questions and no strawman: we gave the chat side more than a real user would.
Run everything on the clean store first
Any test that fails where nothing is wrong is a false alarm, and it is disqualified from the bug hunt. A suite that cries wolf on a working app has not earned the right to be believed on a broken one.
The thing a chat window cannot do
A chat model never finds out it was wrong. It writes a test, the conversation ends, and nothing ever reports back. So we built the loop that does: run the test, and when it fails, hand the model the page as the run actually left it — the URL, the accessibility tree, the visible text, the requests that step made — and ask one question. Was the test wrong, or is the app?
Both answers had to be load-bearing. If the test was wrong, rewrite it and run it again. If the app is wrong, return nothing, leave the test alone, and stop — because repairing a test until it goes green is how you delete a real bug. That refusal is the whole design: every other tool in this category advertises “self-healing”, which is the same mechanism with the safety catch removed.
The number now
| Suite | Passes on the working store | False alarms | Bugs caught |
|---|---|---|---|
| Our engine, before | 12 | 6 | 3 |
| Chat, given every page | 11 | 5 | 4 |
| Our engine, with the loop | 15 | 3 | 5 |
Of the six scenarios that failed on the clean store, the loop repaired three into working tests and insisted three were defects — wrongly, since that store had nothing wrong with it. The repaired three then went on to catch two more seeded bugs, which is the part worth understanding: a test that never runs clean can never catch anything. Half our false alarms were not bugs we missed. They were our own tests being wrong about the app, and they were costing us catches.
What we still get wrong
Three false alarms remain, and each one is a human being asked to dismiss something that was never broken. Five of the ten seeded bugs are still standing — case-sensitive search, stacked discount codes, a shipping threshold, a rounding rule, and a remove button that removes the wrong row — and nothing we have built has ever caught them. This run covered three of the four features, and the chat baseline is the number from the earlier run rather than a same-day re-match.
We publish the losing numbers because a testing tool that only reports its wins is asking you to take the most important thing about it on faith. If you want to see the machinery, Test Automation is in closed beta, and we did this to our matching engine too.