← All posts

Sep 11, 2026

FNOL → Adjudication → Payout Saga

How it works

The obvious places to test are the ends: did FNOL Intake capture the claim, and did the Payment API move the money. Both are cheap to verify and both are misleading, because neither is where the claim's meaning is decided or where it can be silently lost. The decision that matters is made across the seam between Coverage Verification and the Adjudication Rules Engine — the rules engine does not re-derive coverage, it consumes what the verification API hands it, and so every rules outcome is only as correct as the coverage payload it was given. A verification response that is stale, partial, or ambiguous still produces a confident adjudication result, and that result lands in the Approved Payout Queue looking exactly like a correct one. A QA engineer who only checks that intake succeeded and that payment returned 200 has tested the two points in the chain least capable of being wrong in an interesting way.

The second concentration of complexity is the stretch from Approved Payout Queue through Payout Adapter to Payment API, Payment Ledger, and finally the Settlement Batch Job — because this is where the flow stops being a request and becomes a saga. The queue means adjudication and payout are decoupled in time, so the adapter can retry, double-submit, or fail after the Payment API has already acted; the Payment Ledger is the only record that can adjudicate between those possibilities, and the Settlement Batch Job runs later still, reconciling in bulk against state that may have moved. The interesting test cases live in that asynchrony: a payout approved but never adapted, adapted twice but ledgered once, ledgered but missed by settlement. None of these are visible at either endpoint — FNOL looks clean and Payment API looks successful in every one of them. Attention belongs at the queue-to-ledger boundary first, because that is the only region of this flow where a claim can be simultaneously approved, paid, and unaccounted for.

Caveats — what breaks in practice

The failure most likely to blindside a team is the silent gap: a change saved in the FNOL service without an integration event ever being emitted. Every other entry on this list announces itself somehow — duplicates show up twice, poison messages halt processing, DLQ accumulation at least piles up somewhere, retry storms and throttling show up as load, timeouts and validation rejections return errors. A silent gap produces nothing: no error, no retry, no queue depth, no rejected record. Adjudication simply never hears about a claim that exists, and payout never happens, and the only signal is a customer or a reconciliation report weeks later. The assumption that makes it surprising is that persisting a change and publishing its event are one atomic act — that if the write succeeded, the message went out. They are two operations, and the write can commit while the publish does not. The same assumption failing in the other direction is the "success response but record not actually persisted" case, and it is surprising for the same reason: the acknowledgment and the durable state are decoupled, but everyone reads the acknowledgment as proof of the state.

For testing, this means happy-path end-to-end runs through FNOL → Adjudication → Payout will never catch it, because in a healthy run the event does get emitted and the chain completes. You have to test the seam directly: assert that every committed state change in the source has a corresponding emitted event, not that the downstream eventually shows the right answer. That means reconciliation-style checks — count and match source records against emitted events against staged records — rather than transaction-following tests, and it means treating a 200 from the persist call as unverified until the record is read back. It also reframes what a passing pipeline test means: absence of downstream effect is the expected observation when the event was never sent, so a test that only watches the consumer side cannot distinguish "nothing happened because nothing should have" from "nothing happened because the event vanished."

How to test this end to end

Picture FNOL-2024-08817 arriving at the FNOL Intake Service from the mobile app at 14:02:11 UTC: policyNumber "AUTO-4471902", lossDate "2024-08-14", lossType "COLLISION", reportedBy "insured", estimatedAmount 4250.00, and a client-generated idempotencyKey "mob-9f31c7ae". Intake normalizes the app's freeform payload into the canonical claim record — mapping the app's `incident_type` string to the enumerated `lossType`, stamping `receivedAt`, and assigning the claim id — persists it, and emits a `ClaimReported` integration event carrying claimId, policyNumber, lossDate, and lossType. That event is the only thing that wakes up coverage verification, which then calls the policy system live to confirm AUTO-4471902 was in force on 2024-08-14; nothing further in the saga — adjudication, the approve/deny/appeal routing, the authorized-then-batch-settled payout — happens if that event never leaves.

The three ways this goes wrong all bite downstream. If the mobile client times out and retries with the same idempotencyKey "mob-9f31c7ae" but Intake keys only on claim id, you get two `ClaimReported` events and two independent trips through adjudication for one collision — and since approved payment is authorized immediately and settles later in a batch, a duplicate can become two real disbursements before anyone reconciles. If the persist succeeds but the publish fails (no outbox, no transactional guarantee), FNOL-2024-08817 sits in the claims store looking perfectly filed while coverage verification never runs — a silent gap the insured only notices when nothing happens. And if Intake's own model changes — say `estimatedAmount` becomes a nested `{amount, currency}` — consumers parsing the old flat field may drop or misread it. To test: submit FNOL-2024-08817 twice with the same idempotencyKey and assert exactly one `ClaimReported` on the topic and exactly one adjudication outcome; inject a broker failure right after the DB commit and assert the event is eventually published (or the claim is flagged unpublished) rather than lost; and run a contract test pinning the `ClaimReported` schema for claim FNOL-2024-08817 so that renaming or nesting `estimatedAmount` fails the build instead of silently degrading coverage verification.

Picking up FNOL-2024-08817 where CP1 left it: the `ClaimReported` event lands, and Coverage Verification calls the policy system live for AUTO-4471902 as of lossDate 2024-08-14, returning a verification result — say `coverageStatus: "IN_FORCE"`, `deductible: 500.00`, `policyEffective: "2024-01-01"`, `policyExpires: "2025-01-01"` — which is handed to the Adjudication Rules Engine along with lossType "COLLISION" and estimatedAmount 4250.00. The engine applies the deductible, lands on `decision: "APPROVED"` with `payableAmount: 3750.00`, and routes that onto the Approved Payout Queue as message `pay-req-8817-01`, separate from the deny and appeal lanes so an appeal-routing bug can't hold FNOL-2024-08817's money. The Payout Adapter consumes `pay-req-8817-01`, authorizes against the payment system immediately, and marks the claim awaiting the scheduled settlement batch — so the authorization exists before the cash moves.

What bites here is that every hop is a queue with redelivery. If the Payout Adapter takes longer than the consumer visibility timeout to get an authorization back, the broker redelivers `pay-req-8817-01` and a second consumer authorizes 3750.00 again — two authorizations, and because settlement is a deferred batch, both can post before reconciliation catches them. Two messages touching FNOL-2024-08817 concurrently (a late coverage re-check and the payout request) can race on the same claim row and leave the decision and payout state disagreeing. A null or unexpected enum from the policy system — `coverageStatus: "UNKNOWN"` instead of IN_FORCE/LAPSED — can either be mapped to the wrong lane or become a poison message that stalls the queue behind it, and if it drops to the DLQ unalerted, FNOL-2024-08817 just never pays. Auth/session expiry on the payment target turns into a retry storm that throttles the whole approved lane. To test: hold the Payout Adapter's response past the visibility timeout for `pay-req-8817-01`, force redelivery, and assert exactly one authorization for 3750.00 exists (idempotency keyed on the payout request id, not the claim id); replay `pay-req-8817-01` and a coverage re-check concurrently against FNOL-2024-08817 and assert the final decision/payout state is consistent; feed the adjudication engine a coverage response with `coverageStatus: null` and then `"UNKNOWN"` and assert FNOL-2024-08817 routes to a defined lane and does not block the queue; expire the payment system session mid-authorization and assert bounded retries plus a DLQ alert naming claim FNOL-2024-08817 rather than silent accumulation.

Picking up where CP2 handed off: the Payout Adapter's authorization for `pay-req-8817-01` has been accepted by the Payment API, and now that authorization has to land in the Payment Ledger as a real, settleable row before the scheduled settlement batch runs. The call carries the fields the batch will later act on — `claimRef: "FNOL-2024-08817"`, `payoutRequestId: "pay-req-8817-01"`, `amount: 3750.00`, `currency: "USD"`, `payeeAccount: "****4471"`, and a ledger state that starts at `STAGED` and is expected to move to `AUTHORIZED` once the Payment API persists it. The API returns `202 Accepted` with `ledgerEntryId: "LGR-99312"`, and that's the receipt the claim record hangs its "awaiting settlement" status on. Everything downstream — the batch that actually moves the 3750.00 — reads the ledger, not the queue, so whatever is written here is the last word on what the customer gets paid.

The trouble is that a `202` is a promise, not a persistence guarantee. LGR-99312 can be acknowledged and never actually written, or written and left in `STAGED` forever because nothing ever transitions it to `AUTHORIZED`, and since settlement is a deferred batch nobody notices FNOL-2024-08817 didn't pay until the customer calls. Under batch load the Payment API throttles, and the retry that follows a throttled-but-actually-persisted call stages a second entry for `pay-req-8817-01` — two ledger rows, 7500.00 exposure, discovered only at reconciliation. Edge-case values get rejected quietly: a payout that adjudicates to `3750.00` is fine, but a full-deductible loss landing on `payableAmount: 0.00` or a currency the ledger doesn't recognize can bounce validation with an error the adapter treats as retryable. And target-side defaults can silently rewrite the record — the ledger stamping its own `currency: "USD"` or defaulting a missing `effectiveDate` to today, so the amount is right but the posting date disagrees with the claim's lossDate of 2024-08-14. To test: post `pay-req-8817-01`, take the `202` with `ledgerEntryId: "LGR-99312"`, then read the ledger back independently and assert one row exists in `AUTHORIZED`, not `STAGED`, and assert a staged-age alarm fires if it sits past the settlement window; replay the same payout request under induced throttling and assert exactly one ledger row for `pay-req-8817-01` regardless of how many 429s came back; submit FNOL-2024-08817 with `amount: 0.00` and with an unrecognized currency and assert the rejection is classified non-retryable and surfaces named to an operator; and submit the record with `currency` and `effectiveDate` omitted, then diff the persisted LGR-99312 against what was sent to prove no target-side default silently changed the payout.

The nightly settlement batch picks up LGR-99312 among roughly 4,200 `AUTHORIZED` ledger rows, groups them into a journal, and posts the 3750.00 to `payeeAccount: "****4471"` — flipping the row to `SETTLED` with a `settlementBatchId: "STL-2024-08-15-02"` and a posting timestamp, which is the event the claim record for FNOL-2024-08817 is waiting on to move off "awaiting settlement." Nothing about the earlier `202` matters here; the batch reads the ledger, and the 3750.00 either leaves on this run or waits for the next one. Because the batch handles the whole journal as a unit of work, LGR-99312's fate is tied to the other 4,199 rows it happens to be batched with.

That coupling is where it breaks: a timeout at row 2,800 can leave the journal half-posted, with LGR-99312 either settled-but-not-marked or marked-but-not-settled depending on which side committed first, and a job that dies outright leaves it sitting `AUTHORIZED` forever with no retry owner — indistinguishable, from the claim's view, from a batch that simply hasn't run yet, so the customer's SLA quietly expires. Posting order is the subtler one: if LGR-99312 and a later adjustment against the same `payeeAccount` post out of event order, the running balance is wrong even though both rows exist. To test, seed LGR-99312 into a batch of 4,200 and kill the job mid-journal, then assert on restart that LGR-99312 is either fully `SETTLED` once or still `AUTHORIZED` — never both-and-neither — and that rerunning STL-2024-08-15-02 doesn't double-pay the 3750.00; separately, assert an alarm fires when LGR-99312 sits `AUTHORIZED` past its settlement window rather than just waiting for the next run; and post LGR-99312 alongside a later same-account adjustment deliberately delivered first, then assert the resulting balance matches the event order, not the arrival order.

CP1 — Capture & Transform

FNOL Intake Service

CP2 — Event Routing & Execution

Coverage Verification API → Adjudication Rules Engine → Approved Payout Queue → Payout Adapter

CP3 — Target Delivery

Payment API → Payment Ledger

CP4 — Target Posting & Application

Settlement Batch Job

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.