Field notes · Testing

All your tests passed. The bug shipped anyway.

Why test suites are getting worse at catching the bugs that matter, why it’s happening faster every month, and what to do about it.

Picture a Monday morning.

Over the weekend, the platform your online store runs on pushed a routine update. At 7am your regression suite runs. 412 tests. All green. You close the report and get on with your week.

On Wednesday, customer support asks why a few hundred orders have no email address on them. Nobody can send those customers a receipt, a shipping update or a refund.

The update had made the email field at checkout optional. Your tests didn’t notice, because nothing broke. The checkout still worked. It just stopped checking.

If you’ve worked in testing for a while, you’ve probably seen a version of this. The rest of this post is about why it happens, why it’s about to happen a lot more, and what actually helps.

Tests notice what breaks. Not what disappears.

A test checks what you told it to check, on the day you wrote it. That makes it very good at one kind of change and blind to another.

Rename a button and the test fails. Loudly. You spend twenty minutes fixing it, and nothing was ever wrong. Remove a rule, like “email is required”, and the test passes, because the steps it follows still work. The harmless change gets all the attention. The dangerous one gets none.

What changed
“Place order” is relabelled “Complete purchase”Recorded suiteFailsOur diffLow · cosmetic
The layout moves “Add to cart” into a new panelRecorded suiteFailsOur diffLow · cosmetic
Email is no longer required at checkoutRecorded suitePassesOur diffHigh · stopped enforcing
ZIP stops rejecting anything that isn’t five digitsRecorded suitePassesOur diffHigh · stopped enforcing
The cart now says free shipping starts at $75, not $50Recorded suitePassesOur diffRule changed
A new required Phone field appearsRecorded suiteFailsOur diffMedium · stricter
“Recorded suite” means one that replays what it clicked and checks what it saw. It is loud about the harmless rows and silent about the dangerous ones.

Many testing tools now offer “self-healing”: when a test fails, it repairs itself. That takes care of the renamed button. It does nothing for the missing rule, because nothing failed. And if the app changed for a bad reason, healing the test is exactly how the bug gets hidden.

Why this is getting worse

Ten years ago, business software changed when you decided to upgrade it. An upgrade was a project. It had a date, a test cycle and somebody who signed it off. Your test suite got updated because the project forced it to.

That world is going away, for two reasons.

1. Your vendors update the app for you

We’ve spent the last few months researching sixteen of the biggest enterprise platforms, from Salesforce and SAP to Shopify, Epic and Guidewire, to build the knowledge our testing engine uses. For this post we asked each one a simple question: after you go live, how often does it change, and who decides when?

Here is one year of each.

How often each of sixteen enterprise systems changes over one year, from continuous to no public scheduleJFMAMJJASONDHubSpotContinuous, phasedDuck CreekEvery two weeksDynamics 365 F&O4 updates + biweekly buildsAdobe CommerceMonthly security patchesShopify PlusNew API every quarterEpicQuarterly versionsManhattan Active WMQuarterly, versionlessSalesforce3 a year, no opt-outGuidewire Cloud3 named releasesSAP S/4HANA CloudFebruary and AugustBlue YonderTwo a year (1H, 2H)Temenos TransactOne major a yearOracle HealthPlatform movingFiservNo public schedulePTC FlexPLMNo public scheduleNedap iD CloudNo public schedule
Release in a month the vendor namesPatch or fix buildCadence published, months not — spacing is oursNo public schedule we could find

A few things jump out.

  • Many change every few weeks, or all the time. HubSpot ships continuously. Duck Creek and Dynamics 365 push updates every two weeks. Manhattan’s warehouse system doesn’t even have version numbers anymore.
  • You mostly can’t say no. Salesforce upgrades are part of the contract. Dynamics 365 lets you pause one big update at a time, but not the two-weekly ones underneath.
  • Two customers on the same product may not have the same app. Duck Creek ships features switched off and lets each insurer turn them on. HubSpot rolls features out by plan and stage. “Which version are we on?” no longer has one answer.
  • Four of the sixteen don’t publish a schedule at all, at least not anywhere we could find. You can’t plan a test cycle around a date you can’t see.

And that’s before your own team changes anything. Every one of these platforms gets configured heavily for each customer: fields, workflows, approval rules. That changes on your schedule, on top of the vendor’s.

Under the hoodAll sixteen systems, with sources
SystemHow it changesWhat the customer controls
SalesforceThree releases a year: spring, summer, winter.No opt-out. A preview sandbox arrives a few weeks early.
SAP S/4HANA Cloud, public editionTwo upgrades a year, in February and August.Rolled out on SAP’s calendar, system by system.
Dynamics 365 Finance & OperationsFour service updates a year, plus fix builds every two weeks.You can pause one service update at a time. The two-weekly builds can’t be paused.
Shopify PlusA new API version every quarter.Each version is supported for at least 12 months, then you have to move.
Adobe CommerceSecurity patches every month except May, and a full patch each May.Yours to apply, until a Cloud upgrade deadline says otherwise.
HubSpotContinuous, with monthly roundups.Features reach accounts at different times, by plan, beta opt-in and rollout stage.
Guidewire CloudThree named releases a year.Not stated on the public release page.
Duck CreekUpdates every two weeks, applied automatically.New features arrive switched off; the insurer decides when to turn them on.
Temenos Transact*One major release a year.Banks are encouraged to upgrade at least yearly; on SaaS, Temenos runs it.
Epic*A new version every quarter, and a move from a desktop app to a browser-based one.Epic suggests upgrading every 12 to 18 months; safety fixes arrive in between.
Oracle Health*Moving customers to Oracle’s cloud and toward a new EHR.No public release schedule found.
Manhattan Active WMNew features every quarter, no version numbers, no downtime.Sold as “never upgrade again”, which also means never choosing when.
Blue YonderTwo named releases a year.Not stated publicly.
FiservNo public schedule found.—
PTC FlexPLMNo public schedule found.—
Nedap iD CloudNo public schedule found.—

Each name links to where the row comes from. * means we could only find a secondary source.

2. More of the code is written by AI

In October 2024, Google said about a quarter of its new code was written by AI. In April 2026 it said this: “75% of all new code at Google is now AI-generated and approved by engineers.” That’s eighteen months. Your vendors are software companies too, and they’re using the same tools.

More code means more change, and more change is where things slip. Google’s own DORA research, which surveyed nearly 5,000 engineers in 2025, found that 90% now use AI at work and that teams are shipping faster. It also found stability going the other way, and it named the fix: strong automated testing and fast feedback. In other words, the part of the pipeline that’s expected to catch all this extra change is testing. Which is the part that still works the way it did ten years ago.

Put the two together, and this is what a test suite looks like over one quarter:

Over twelve weeks the app changes fourteen times while the test suite still describes week zeroweek 4week 8week 12tests writtenthe app todaywhat the tests still believedrift
Vendor releaseConfiguration changeMerged pull request (often agent-written)Dependency bumpAn illustration, not a measurement. The dashed line never moves unless someone moves it.

What actually helps

None of this needs a particular tool. These are the habits that close most of the gap, and you can start on them this week.

1

Test that rules are enforced, not just that screens work

Most suites check the happy path: fill the form, submit, see the confirmation. Add the unhappy one. Submit without an email and expect to be stopped. Enter a ZIP code of “abc” and expect an error. Those are the tests that notice when a rule quietly goes away.

2

Check after every change, not every release

If the app changes every two weeks, a quarterly regression cycle is testing a version that no longer exists. Hook your checks to whatever knows something changed: a deploy, a vendor’s release email, a merged pull request. When nothing tells you, run them every day.

3

Compare today’s app with yesterday’s

Keep a record of what the app looked like: its pages, its fields, what each field demanded. When something changes, the difference tells you exactly where to look, before any test has a chance to fail or, worse, pass.

4

Before you fix a failing test, ask who’s wrong

A failing test means one of two things: the test is out of date, or the app has a bug. Auto-fixing tests assumes it’s always the first. Make someone, or something, answer the question before the test gets changed.

How we built this into QualityAIQ

Those four habits are simple to say and tiring to do by hand every day. So we built a loop that does them. The key decision was what to save. Most tools save tests. We save the app itself: what it looked like, what each field demanded, and what each page said, every time we looked. Tests are made from that, and can be made again when the app moves.

  1. 1

    Something moved

    A deploy hook, CI, a merged PR, a vendor notice, or a timer.

  2. 2

    Explore

    Walk the app in a real browser and store a snapshot.

  3. 3

    Read

    Pull the rules the pages state in words, quoted exactly.

  4. 4

    Diff

    Compare with the last snapshot. Grade what got weaker.

  5. 5

    Map

    Find the tests that touch what changed, and how closely.

  6. 6

    Run

    Only those, most dangerous first.

  7. 7

    Judge

    For each failure: is the test wrong, or the app?

  8. 8

    Report

    Clear, defects, needs a person, or not done yet.

And underneath, a ninth step: every time a person confirms or corrects step 7, it counts against us. That’s how we find out whether our confidence is worth anything.

It looks at the app the way a person would

Our explorer opens the app in a real browser and clicks around, because a cart with something in it is a different page from an empty one. It writes down every form, every field, and what each field requires. Each visit is saved, so “what changed?” is just a comparison with the last one.

It reads the rules written on the page

Some of the most important rules aren’t in the code you can see. They’re in a sentence: “One code per order.” “Free shipping on orders of $50 or more.” We use AI to pull those sentences out and turn them into checks, and we throw away any rule that isn’t quoted word for word from the page. A made-up rule becomes a test that fails on a perfectly good app, and that’s worse than no test. If free shipping moves from $50 to $75, nothing in the page’s code changes at all. We still catch it, because the sentence changed.

Under the hoodHow strict the quoting is, and what the first run found

The model reads each page’s text once per visit, never once per test run, so the cost doesn’t grow with the suite. Every proposed rule has to carry a quote that appears on that page exactly. For our knowledge base we require quotes of at least 25 characters so a short fragment can’t match by accident, but on product pages that rejected the best rules: “One code per order.” is 19 characters. Page quotes use a 16-character minimum.

On our test store, the first run proposed six rules and kept all six. The pages that state no business rules, the home page and checkout, came back with none rather than invented ones. A reworded rule is reported as one change with both quotes, not as a removal plus an addition.

It knows what matters in your industry

A page tells you what the app says. It doesn’t tell you what the business expects: that a purchase order has to match its invoice, or that an insurance claim can’t close while money is still reserved against it. That’s what the research on those sixteen platforms is for: 160 detailed write-ups of how their modules work. It lets us tell a harmless change on a product page from a serious one in the finance module. This part is still early. We’ve turned one of the 160 into checkable rules so far.

It ranks changes by what got weaker

This is the opposite of how most tools think. A rule that stopped being enforced is ranked above a broken button, because the broken button will announce itself and the missing rule never will.

High

The app stopped enforcing something it used to: a required flag dropped, a pattern removed, a limit raised, a field or page gone, a page that now loads with a failing request.

Medium

The app got stricter, or grew something nobody tests yet: a tighter limit, a new required field, a new form.

Low

Cosmetic: a new title, a new button, a different status code that still succeeds.

It re-runs only what the change touched

On our test store, we removed the rule that a ZIP code must be five digits, the kind of thing a quiet update might do. Then we let the loop run. It re-ran 9 tests in 31 seconds, not the whole suite. Eight passed. The one that failed was the test that types “abc” into the ZIP field and expects an error, because the app had genuinely stopped giving one. It’s the same kind of change as the Monday-morning email bug, caught on Monday.

Under the hoodOur first two attempts were too noisy, and how we fixed them

Run one flagged every checkout test for every checkout change, because every one of them starts by opening the checkout page. We split matches into two strengths: a test step that touches the exact thing that changed, and a test that merely passes through the page. The second still gets re-run, but can never be blamed on the changed rule.

Run two still said seven of nine checks were built on a rule that had moved, when only two rules had moved. The cause: to test the Name field, a check fills in every other field so the form is otherwise valid, which means it legitimately types into Email too. Now each check is matched against its own field only, by exact id. Result: four flagged, all four correct.

Pages are matched by address and by how they were reached, so an empty cart is compared with the last empty cart, not with a full one.

It asks who’s wrong before fixing anything

When a test fails, the loop asks one question: is the test out of date, or is the app broken? If the test is out of date, it’s rewritten against the page as it really is. If the app is broken, the test is left alone and the failure is reported. And when it can’t tell, it says “needs a person” instead of guessing. A cautious answer is more useful than a confident wrong one.

It runs whenever something changes

Since four of those sixteen platforms don’t publish a schedule, the loop can’t wait for one. It can be started by anything that knows the app moved: a deploy, a vendor’s release email, a merged pull request, or a daily timer.

Under the hoodThe webhook, and what it’s allowed to do
curl -X POST https://www.qualityaiq.com/api/automation/hooks/<app-id> \
     -H 'authorization: Bearer <token>' \
     -d '{"ref":"deploy 4c1f2ab"}'

The token can only re-check the one application it belongs to, is stored only as a hash, and a wrong token looks exactly like an unknown application. Every call leaves a record of what changed and what was done about it, which is usually the first thing someone asks for after a bad release.

It keeps score on itself

Every “test or app?” answer comes with a confidence level. When a person confirms or corrects it, we count it. If we turn out to be more confident than we are right, the product says so, in those words: “Overconfident: we claim X and are right Y of the time.” With fewer than ten answers checked, it refuses to give a number at all. A tool that wants a place in your release process should be able to tell you when not to trust it.

Under the hoodHow the scoring works

We use a Brier score, which punishes a confident mistake much harder than a hedged one, and report the gap between stated confidence and observed accuracy per confidence band. Rules and model answers are scored separately, because they fail in different ways. It needs no model call and nothing to train; it reads the labels already in the database.

What we haven’t proven yet

  • A real vendor update. Everything above was tested on a store we built, where we made the changes ourselves. Next is a real Salesforce org through a real seasonal release, where we don’t know what’s coming.
  • Rules turning into catches. Reading rules off the page works. We haven’t yet measured how many extra bugs that catches.
  • The industry knowledge. One of 160 write-ups has been turned into rules.
  • The self-scoring. It works, but it needs real people checking real failures before its number means anything.

The question has changed

For years the question was: did the upgrade break anything? That question assumed there was an upgrade. More and more, there isn’t. There’s just an app that’s a little different from yesterday.

The better question is: what changed since we last looked, and which of our checks does it touch? You can ask that every day, cheaply, if you keep track of the app and not just the tests. That’s what Test Automation is built to do. And if you’d like to see how we measure ourselves, here’s the benchmark we published while we were losing.

Sources

  1. Google, “Sundar Pichai shares news from Google Cloud Next 2026”
  2. Fortune, “Google CEO: more than 25% of new code is AI-generated” (Oct 2024)
  3. Google Cloud, “Announcing the 2025 DORA Report”
  4. GitClear, “AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones”
  5. Microsoft Learn, “Pause service updates through Lifecycle Services”
  6. Salesforce Help, “Sandbox Preview Instructions”
  7. Whatfix, “Epic Hyperdrive Migration Plan for Healthcare IT Teams”
  8. Salesforce: how it’s released
  9. SAP S/4HANA Cloud, public edition: how it’s released
  10. Dynamics 365 Finance & Operations: how it’s released
  11. Shopify Plus: how it’s released
  12. Adobe Commerce: how it’s released
  13. HubSpot: how it’s released
  14. Guidewire Cloud: how it’s released
  15. Duck Creek: how it’s released
  16. Temenos Transact: how it’s released
  17. Epic: how it’s released
  18. Oracle Health: how it’s released
  19. Manhattan Active WM: how it’s released
  20. Blue Yonder: how it’s released