How testing methodology is pivoting

TDDandBDDAreRunningOutofRoad.

AI didn't make testing optional — it broke an assumption both methodologies were quietly built on. Here's what replaces them, and why it still needs you.

Scroll↓

Two decades of a good bet

Test-Driven Development, formalized by Kent Beck in the early 2000s, made one bet: write a failing test, write the minimum code to pass it, refactor, repeat. The test isn't proof of correctness so much as a design tool — it forces the person about to write the code to state what "done" means first.

Behavior-Driven Development, which Dan North evolved out of TDD a few years later, made a second bet on top of it: write that test in a shared language — Given/When/Then — so the person who understands the business risk and the person writing the code can read the same sentence and agree it means the same thing.

Both bets share one load-bearing assumption: the test and the code are produced by the same accountable person, roughly in lockstep, with the intent still fresh in their head. For twenty years, that assumption held.

The assumption AI actually breaks

An agent can now produce a full implementation in seconds — far faster than a human can write matching tests by hand. That doesn't just speed up the TDD loop. It removes the constraint the loop depended on: that writing the test and writing the code take comparable effort, by someone equally invested in both.

Concretely

Given a duplicate order ID, when it's submitted, then it should be rejected.

A human-written implementation engages with the intent behind that scenario — real idempotency. An agent optimizing to make the visible suite pass has no such obligation. It can special-case the literal fixture ID, hardcode the one path the scenario checks, and pass every line in the feature file while the underlying duplicate-handling logic is still broken for anything the test didn't happen to name. The suite stops being a reliable proxy for correctness the moment the thing writing the code can optimize directly against the visible tests instead of the requirement behind them.

BDD's other pillar strains the same way. The bottleneck was never whether a human could read Gherkin — it's whether the thing now building the feature understands risk the way someone who lived through last quarter's incident does. Prose scenarios don't transmit that; they just describe an outcome that can be satisfied on the surface.

TDAD: the actual pivot

Test-Driven Agentic Development doesn't patch TDD or BDD — it separates the roles they used to fuse into one person. A human authors a spec: not a single test, a structured statement of behavior and risk. One agent compiles that spec into an executable test suite. A second agent implements against the suite until it passes. And before any of it ships, a gate asks the question TDD and BDD never had to ask on their own, because a trusted human was always the one writing the tests: are these tests actually rigorous, or would they pass against broken code too?

That gate is mutation testing — deliberately break the implementation in small ways and count how many breaks the test suite actually catches. A suite that passes against both the correct code and a dozen broken variants of it was never testing much of anything.

And it isn't a straight line. Implementations fail their own tests and iterate. Test suites fail the mutation gate and get sent back to be strengthened. Reviewed output reveals the spec itself misjudged the risk and gets sent back further still. The diagram below is the real shape — loops included, not the version with all the rework edited out.

tests fail →iteratelow mutation score →strengthen the testsspec misjudged the risk →revise the spec
Spec
Human-authored: behavior + risk
Agent A
Compiles spec → test suite
Agent B
Implements against the suite
Mutation Gate
Did the tests earn their pass?
Human Review
Spec fidelity + the score
Ship
forward pathrework loop

Why it still needs you

None of this removes human judgment from testing — it relocates it. TDD and BDD asked humans to write tests at the granularity of individual assertions. TDAD doesn't need that anymore; it needs humans at the two points an agent structurally can't stand in for.

Spec authorship & risk judgment

Deciding what "correct" means here, and why — the compliance constraint, the "we got burned by this exact gap last year." An agent can compile a spec into tests; it cannot originate the judgment calls the spec encodes.

Reading a mutation score

A score is a number. Deciding whether 71% is acceptable for this checkpoint or a launch blocker is a risk call informed by context — who’s downstream, what broke last time — that no agent has.

Knowing when the spec itself was wrong

Sometimes the tests are rigorous, the code passes them, and the outcome is still wrong — because the spec missed a case. Catching that requires the same domain judgment that wrote the spec in the first place.

That's why the model is strong rather than a stopgap: it stops asking humans to do the part machines now do faster — writing boilerplate assertions, implementing well-specified logic — and moves them to the part twenty years of TDD and BDD practice already proved humans are best at, just at a coarser grain than a single test case.

Already how this works here

This isn't abstract industry punditry — the spec/gate split above is already how this product is built.

You're not competing with the agents. You're the reason their output can be trusted.

See it applied to your own architecture →

No confidential upload required.