← All posts

Sep 6, 2026

Subscription Billing Dunning

How it works

The obvious places to look in this flow are the two endpoints: does the Billing Engine produce the right charge amount, and does the Payment Gateway return the right decline code. Both are worth a handful of tests and neither is where subscription billing actually breaks. The real complexity sits in the Dunning Retry Job, because it is the only component in the chain that has to hold state across time. The Billing Engine fires once and is done. The Gateway answers once and is done. The Retry Job must remember that a specific charge failed, decide when to try it again, decide how many times it is willing to try, and decide when to stop — and it must make all of those decisions correctly while the Billing Engine keeps generating new charges for the same subscription on its own schedule. That overlap is not an edge case; it is the normal condition of any account that stays delinquent longer than one billing period.

For a QA engineer this means the highest-value tests are not "valid card succeeds" and "invalid card declines," they are tests about what the Retry Job does with a failure it has already seen. What happens when a retry succeeds after the Billing Engine has already issued the next cycle's charge — does the customer get charged twice, and does the dunning state actually clear? What happens when a retry attempt itself fails at the Gateway: does the job re-queue it as a fresh failure, restarting the retry count, or does it correctly recognize this as attempt three of three? What happens when the Gateway returns a decline that is not a hard failure — does the job distinguish that from a dead card, or treat all non-success identically and burn its retry budget in an afternoon? None of these questions can be answered by testing the Billing Engine or the Gateway in isolation, because none of them are about a single transaction. They are about the Retry Job's memory, and memory is where this flow keeps its bugs.

Caveats — what breaks in practice

The failure mode most likely to blindside a team is the one where a successful retry doesn't cancel the remaining scheduled attempts and the customer gets double-charged. Every other item on the list announces itself in a way teams already instrument for: throttling under batch load shows up as latency and error rates, validation rejections on edge-case values leave rejected records, schema drift breaks consumers loudly, and duplicate events on retry after a transient failure are the textbook case everyone writes idempotency keys for. The uncancelled schedule is different because nothing fails. The charge succeeds, the dunning sequence is technically doing what it was told, and the only signal is a second successful charge that looks identical to a legitimate one — until the customer complains.

The assumption doing the damage is that success is terminal: that recovering the payment retires the dunning state machine and the attempts still sitting on the schedule become moot. It's a reasonable expectation, because the whole point of a retry schedule is to stop when it gets what it wanted. But the schedule is a separate mechanism from the outcome, and the same decoupling shows up twice more in this list — a retry schedule that skips a step under load and suspends a customer early, and suspension applied before the grace period's final retry has actually run. All three are the same bug wearing different clothes: timeline state and payment state drifting apart. For testing, this means the interesting assertions are not "did the charge succeed" but "what is left on the schedule after it succeeded," which requires inspecting scheduled future work as an observable, not just the response to the call you made. Pair that with the silent-gap mode — changes saved without emitting an integration event — and you have a system where the scheduler's view of a customer can be stale in both directions with no error anywhere to trace.

How to test this end to end

Picture invoice `inv_9f42c1` for subscription `sub_88231` on the Pro Annual plan, $480.00 USD, charged on renewal day and declined by the processor with `insufficient_funds`. Inside the Billing Engine at CP1, that decline is captured and transformed into a dunning state record — `dunning_state: "retry_scheduled"`, `attempt_number: 1`, `next_attempt_at: 2024-06-03T09:00:00Z`, `grace_period_ends_at: 2024-06-14T09:00:00Z`, `failure_code: "insufficient_funds"` — and the engine emits an integration event, say `invoice.payment_failed.v2`, carrying that same payload so the scheduled retry sequence and escalating notifications downstream have something to act on. The critical transform here is that a single processor decline becomes a *durable multi-day plan*, not a one-shot outcome: the record on `inv_9f42c1` is what every later attempt and the eventual suspension decision reads from.

The failure modes bite in three shapes. The engine can write `dunning_state: "retry_scheduled"` on `inv_9f42c1` and never emit `invoice.payment_failed.v2` — a silent gap where the DB says dunning is running but nothing downstream is scheduled, so the customer is never notified and never suspended. Or a transient processor timeout causes the capture to be retried and two `invoice.payment_failed.v2` events land for the same `inv_9f42c1` / `attempt_number: 1`, which duplicates the scheduled attempts and sets up the double-charge that the mid-sequence cancellation is supposed to prevent. Or the engine's own model changes — `failure_code` becomes a nested `failure: {code, network_decline_code}` — and the event schema drifts out from under consumers that still read a flat string. Test it by forcing a declined charge on `inv_9f42c1` and asserting both sides in one go: the persisted dunning row and exactly one emitted event, matched on `(invoice_id, attempt_number)`, with a schema assertion against the published `v2` contract. Then replay the same capture with an injected transient failure and confirm `inv_9f42c1` still yields exactly one event, not two. Finally, let attempt 2 on `inv_9f42c1` succeed and assert that no further attempts remain scheduled for that invoice — the cancellation contract, verified on the same record rather than on a fresh one.

At CP2 the plan written at CP1 stops being a record and starts spending money: at `2024-06-03T09:00:00Z` the Dunning Retry Job wakes for `inv_9f42c1`, reads `attempt_number: 1`, `next_attempt_at`, and `grace_period_ends_at: 2024-06-14T09:00:00Z`, and re-presents the $480.00 USD charge on `sub_88231` to the Payment Gateway. The gateway answers — say a second `insufficient_funds` — and that answer has to be posted back and *applied* to the same dunning record: `attempt_number` advances to 2, `next_attempt_at` moves to the next slot (`2024-06-08T09:00:00Z`), `dunning_state` stays `"retry_scheduled"` because `grace_period_ends_at` has not passed, and the escalated notification for attempt 2 goes out. The transform is the mirror of CP1: a single gateway response becomes an update to the durable multi-day plan, and every subsequent attempt plus the eventual suspension decision reads the row this checkpoint just wrote.

Things go wrong in the gap between "the gateway said something" and "the row on `inv_9f42c1` reflects it." The gateway can return a success-shaped response while the write is lost, so the job re-presents $480.00 against `sub_88231` on the next tick as if attempt 2 never happened. Under batch load — renewal day for thousands of subscriptions — the job can be throttled and skip the `2024-06-08T09:00:00Z` slot entirely, then evaluate `grace_period_ends_at` and suspend `sub_88231` before the grace period's final retry actually ran. Validation can reject the posting on an edge-case value (a gateway returning an empty or unmapped `failure_code` where CP1 wrote a clean `"insufficient_funds"`), stranding `inv_9f42c1` mid-sequence. And the cancellation contract still bites here: if attempt 2 *succeeds*, the remaining slot must be torn down or `sub_88231` gets charged $480.00 twice. Test it by driving `inv_9f42c1` through the scheduler with a faked clock: run the `2024-06-03T09:00:00Z` attempt, force a decline, and assert the persisted row reads `attempt_number: 2` and `next_attempt_at: 2024-06-08T09:00:00Z` with `dunning_state` still `"retry_scheduled"` — then re-read it from the database rather than trusting the job's return value, which catches the phantom-success case. Run the same attempt with the gateway returning a blank `failure_code` and assert `inv_9f42c1` still advances rather than being rejected. Queue several hundred synthetic invoices alongside `inv_9f42c1` to trigger throttling and assert `inv_9f42c1` never reaches a suspended state while its clock is before `2024-06-14T09:00:00Z` and any scheduled attempt is still pending. Finally let attempt 2 on `inv_9f42c1` succeed and assert zero remaining scheduled attempts and exactly one successful capture of $480.00 against `sub_88231`.

CP1 — Capture & Transform

Billing Engine

CP2 — Target Posting & Application

Payment Gateway → Dunning Retry Job

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.