Bringing machine learning into how we test

WeMutation-TestedOurOwnMatchingEngine.

We hand you a confidence score every time the wizard matches your architecture. This is our first machine learning model — trained, evaluated, and mutation-tested with the same rigor we tell you to apply to your own systems, before we let it anywhere near what you see.

Scroll↓

5,800

simulated wizard sessions, 3 random seeds

11

reference patterns found to be indistinguishable

10.4% → 2.4%

confidently-wrong rate, before vs. after

0

real user data used to get there

Why we're bringing machine learning into QualityAIQ

QualityAIQ's core has always been deterministic — checkpoints, failure modes, and risk scores computed from data, not guessed by a model. That's not changing. But some parts of testing are genuinely pattern-matching problems, and pattern-matching is exactly what machine learning is built for. Deciding which reference architecture you actually meant, from a handful of partial answers, is one of those problems.

This is our first machine learning model in production-facing work, and we're starting it on ourselves on purpose: the matching engine, not your architecture. If we're going to adopt ML into how QualityAIQ does quality engineering, the bar is that it gets the same testing discipline we'd demand from anyone else's AI component — measured against a baseline, verified across seeds, reproduced in the cloud, published either way.

Why turn it on ourselves

The wizard matches your architecture to one of our reference patterns through a short adaptive questionnaire, then reports a confidence score — how sure it is you're looking at the right one. That score comes from a set of weights: which facet matters more, which matters less, hand-assigned when the matcher was built.

Nobody had ever measured those weights against real ambiguity. Every checkpoint we hand you is built on the idea that you test the seams where things quietly go wrong, not just the happy path — so we pointed that same idea at the matcher itself, before shipping it as a claim rather than a measurement.

The experiment

Mutation testing means deliberately breaking something small and checking whether your safety net catches it. We applied it to the matcher itself, not to a diagram:

Give it a real pattern to find

Pick one of our 58 reference architectures and treat it as ground truth — this is the pattern a simulated tester "really" has in mind.

Answer like a real tester would

Walk through our own adaptive questionnaire, but not perfectly: sometimes a facet gets answered correctly, sometimes it gets misjudged, sometimes it gets skipped — the same slips a real person makes.

Run it through the real code

Not a mockup of the matcher — the actual functions the product calls today, unmodified, scoring every candidate the same way a live session would.

Check who it landed on

Did the top pick match the pattern we started with? And critically: was it wrong while still reporting high confidence — the case where a tester gets handed the wrong checklist and no reason to doubt it.

We ran that loop 5,800 times across all 58 patterns, and repeated the whole thing on three separate random seeds so a lucky run couldn't pass for a real result.

The bug we found first, before training anything

Even with every question answered perfectly — zero simulated confusion — the matcher still landed on the wrong pattern 4.2% of the time. That number showed up before we trained a single model, in the plain fuzzing pass that generated the data — worth separating from the machine learning result below, because it's a different kind of finding, fixed a different way.

The actual cause

Five groups of our reference patterns — eleven entries in total — share byte-identical answers across every question the wizard asks. Two genuinely different architectures, answered the same way every time, are the same input to the matcher. No amount of careful answering fixes that; the information needed to tell them apart simply isn't being asked for yet.

That's now a tracked fix, separate from anything below — the honest kind of bug report, the kind we'd want a tool doing this to us to produce.

What actually moved the number

With the fixed ceiling from above accounted for, we trained a machine learning model — logistic regression, deliberately simple and inspectable, not an opaque black box — on which facets got confirmed or contradicted in each simulated session, deliberately withholding the matcher's own confidence score as an input so it couldn't just copy the answer.

Top-1 accuracy barely moved — expected, given eleven patterns are simply unsolvable by facets alone. But the number this whole exercise was aimed at, the dangerous one, moved a lot: the confidently-wrong rate fell from 10.4% to 2.4%, a drop that held steady across all three random seeds, and reproduced byte-for-byte when we re-ran it as a real training job on Google Cloud instead of on a laptop.

The learned weights didn't contradict the hand-tuned ones, either — they refined them. The facet we'd already weighted lowest by hand came out weakest in the learned model too; the manufacturing-specific facet came out stronger than we'd guessed. Measurement agreeing with intuition, in the places intuition was right, is exactly what you'd want to see before trusting the places it corrects you.

What this doesn't touch

Every session in this experiment was simulated against our own reference library — never a real architecture, never a name a real user typed. That's not a footnote; it's the same "never used to train anything" promise on our About page, and this work was built to keep it, not test its edges.

Where this stands

Nothing above is live in the wizard yet — a result worth trusting and a result worth shipping to real users are two different bars, and we're still deciding how a machine-learning-backed confidence score should actually show up for you. This is the first place we're adopting ML into QualityAIQ's own methodology, not the last — and every time we do, we'll report it the same way we just did: the real number, including the parts that didn't move.

The confidence score you see today is still the hand-tuned one — see how it does on your own architecture.

Try the matcher →

No confidential upload required.