← All posts

Sep 4, 2026

Returns Disposition Engine

How it works

The tempting places to aim testing here are the two endpoints: Returns Intake, because it's where a human touches the system and produces obvious defects, and the ERP Inventory & Refund Record, because it's where money and stock counts land and where an error is loudest. But both ends are the most constrained parts of the flow. Intake produces a bounded set of inputs — a return initiated at POS or on the web — and the ERP record is a terminal write whose correctness is determined almost entirely by what arrived. The real complexity lives in the middle, at the seam between the Condition/Barcode Scan, the Disposition Scoring Engine, and the Resell Channel Adapter, because that is the only stretch of the flow where a decision is *made* rather than captured or recorded. The scan supplies condition data of variable fidelity; the scoring engine turns that into a disposition; the adapter translates that disposition into a channel-specific action. Each hop is a translation, and translations are where meaning is lost.

For a QA engineer this means the first-order question is not "did intake accept the return" or "did the ERP balance," but "does the same scanned condition reliably produce the same disposition, and does that disposition survive the adapter intact?" A return can pass intake cleanly, scan cleanly, and still be scored into a disposition the Resell Channel Adapter cannot express, at which point the ERP inventory and refund record is written against a state nobody intended — and it will look internally consistent, because the ERP faithfully records whatever it was handed. That is the failure mode worth designing tests around: not a crash at an endpoint, but a quietly wrong disposition that propagates into inventory and refund with full plausibility. Test the scoring engine's boundary conditions and the adapter's handling of every disposition it can be handed, and the endpoints largely take care of themselves; test the endpoints first and you will confirm that a broken decision was transmitted accurately.

Caveats — what breaks in practice

The failure mode most likely to blindside a team is the model retrain that silently changes scoring thresholds without a corresponding recovery-rate review. Every other item on the list announces itself somehow: a timeout on an oversized batch throws, a poison message halts processing, a retry storm on persistent target failure lights up the target's error rate, an auth/session expiry fails loudly on the next call. Even the quieter ones — duplicate events on retry, records stuck in staged state forever, duplicate staged records from upstream redelivery — leave countable artifacts that a reconciliation check can find. A retrain leaves nothing anomalous behind. The engine keeps accepting events, keeps scoring, keeps routing, and every individual disposition looks structurally valid. What changed is only where the line sits, and the consequence — condition-scoring drift routing genuinely sellable stock to liquidation or disposal — is indistinguishable from a run of genuinely poor returns unless someone is watching recovery rate, which is precisely the review the failure mode says didn't happen.

The assumption that breaks is that the service's behavior is pinned by its deployed code and configuration — that if nobody shipped a change, the disposition logic today matches the disposition logic last week. That assumption holds for the mapping errors, the batch split, the concurrency races; it does not hold for a scored decision whose thresholds come from a retrained model. For QA this means the usual regression posture is aimed at the wrong surface. Contract tests on event schema catch schema drift when the service's own model changes; field-level assertions catch mapping errors on nulls and unexpected enums; neither notices that a correctly-shaped event now carries a different disposition than it would have yesterday. The testable artifact has to be the outcome distribution, not the message: recovery rate per category as a gated check tied to retrain, plus golden-set items with known-correct dispositions replayed after every model change, so a threshold shift fails a test instead of quietly showing up as margin loss.

How to test this end to end

Picture return RMA-4471902 arriving at the Bellevue store counter: the associate scans the customer's receipt, the POS captures `rma_id: "RMA-4471902"`, `order_line_id: "OL-88231-3"`, `sku: "APL-AIRPD-PRO2"`, `return_reason_code: "DOA"`, and `intake_channel: "POS"`, then runs the item across the condition scanner, which appends `barcode_scan: "194253397724"`, `condition_grade: "B"`, `packaging_intact: false`, and `scan_confidence: 0.62`. That combined record is what Capture & Transform is responsible for turning into a single integration event for the disposition engine downstream — everything the model will later use to choose resell, liquidate, or discard originates here, so if a field is dropped or mistyped at this checkpoint, the scoring engine is reasoning about an item it has never accurately seen. The same intake also feeds the refund path, which is exactly why this checkpoint matters: the refund for RMA-4471902 can settle while the disposition event for it is still missing or malformed.

The ways this goes sideways are mundane and expensive. The scan row can be written to the returns table with no event emitted at all — a silent gap where RMA-4471902 exists in the store's records but never reaches the disposition engine, leaving a refunded customer and an item with no recovery channel. A transient scanner timeout followed by a retry can emit RMA-4471902 twice, so one physical pair of earbuds gets scored twice and potentially routed to two channels. And `condition_grade` is the classic mapping hazard: if the scanner returns a null or an unexpected enum like `"UNGRADED"` when `scan_confidence` is low, a naive transform may coerce it to the worst case and send sellable stock toward discard. To test, replay RMA-4471902 through intake and assert exactly one event lands downstream with all eleven fields intact; then force a post-write failure between the scan save and the emit, and confirm the gap is detected rather than swallowed. Retry the same scan three times and assert the disposition engine sees one event keyed on `rma_id` + `order_line_id`, not three. Feed variants of RMA-4471902 with `condition_grade` null, empty, and `"UNGRADED"` and assert the transform rejects or flags them rather than silently defaulting to discard. Submit a batch containing RMA-4471902 alongside 500 other lines and assert line-item counts match in and out, and that one poison record — say a malformed `barcode_scan` — is quarantined without blocking RMA-4471902 behind it. Finally, assert that any refund marked complete for RMA-4471902 has a corresponding disposition event or an open exception, so the refunded-but-undispositioned window is visible rather than invisible.

The event for RMA-4471902 lands in the Disposition Scoring Engine carrying `condition_grade: "B"`, `packaging_intact: false`, `scan_confidence: 0.62`, and `sku: "APL-AIRPD-PRO2"`, and the engine emits a decision — say `disposition: "RESELL"`, `channel_score: 0.71`, `model_version: "cond-score-4.2"`, `target_channel: "RESELL_OPEN_BOX"` — which the Resell Channel Adapter must translate into a call the resell system accepts, mapping `sku` and `condition_grade` onto that system's own listing grade and handling rules. The exposure here is that a correct decision still produces a wrong outcome. Scoring drift or a quiet retrain from `cond-score-4.2` to `4.3` can nudge the threshold so a `B` grade at `scan_confidence: 0.62` now scores 0.48 and RMA-4471902 goes to liquidation instead of open-box resell, with no recovery-rate review to catch it. A misread `barcode_scan` upstream can hand the adapter a SKU whose category has different handling rules than earbuds, so a battery-bearing item enters a channel that doesn't expect one. The adapter itself can fail on field mapping for the open-box subtype specifically, on an expired auth session to the resell API, or by retry-storming a target that is persistently down; and if a second message for the same `rma_id` + `order_line_id` arrives concurrently, two writers can race the same disposition record and leave RMA-4471902 with a decision that doesn't match what was actually executed.

Test it by pinning RMA-4471902 as a golden record and asserting `cond-score-4.2` routes it to `RESELL_OPEN_BOX` at `channel_score: 0.71`, then running the same record against any candidate model version and failing the build if the decision flips without an explicit recovery-rate sign-off. Replay RMA-4471902 with `scan_confidence` swept from 0.45 to 0.95 and chart where the resell/liquidate boundary sits, so drift is measurable rather than anecdotal. Substitute a mismatched `barcode_scan` that resolves to a non-earbud category and assert the adapter rejects the handoff instead of listing it. For the adapter itself, force an auth expiry mid-call and assert RMA-4471902 is re-authenticated and delivered exactly once, not duplicated; force a persistent 503 from the resell API and assert backoff with a capped retry count and a dead-letter entry rather than an unbounded storm. Finally, fire two concurrent decision messages for `RMA-4471902` / `OL-88231-3` and assert the stored disposition and the actual adapter call agree — one winner, one rejected — and that the losing message surfaces as an exception rather than overwriting silently.

Once the Resell Channel Adapter succeeds, the last hop is the ERP, where RMA-4471902 has to land as two linked facts: an inventory record for `APL-AIRPD-PRO2` received back into an open-box location, and a refund record against `OL-88231-3`. The payload arriving at the ERP carries something like `rma_id: "RMA-4471902"`, `order_line_id: "OL-88231-3"`, `disposition: "RESELL"`, `target_channel: "RESELL_OPEN_BOX"`, `condition_grade: "B"`, and a refund amount, and the ERP's returns interface typically stages it before posting. Three things go wrong here. The staged record can simply sit — accepted with a 200, an ERP staging id issued, but never posted, so RMA-4471902 shows as resold in the disposition store and as nothing at all in sellable inventory, and nobody notices because the upstream call succeeded. The ERP can also silently apply its own defaults: `condition_grade: "B"` maps to a listing grade the ERP doesn't recognize and it falls back to `NEW`, or the inventory location defaults to the main sellable bin instead of open-box, so a B-grade earbud set with broken packaging is now priced as new. And because upstream redelivery is a normal thing — the adapter retried, or the decision message was reprocessed — the ERP can end up with two staged records for `RMA-4471902` / `OL-88231-3`, which at post time becomes two units of inventory from one physical return, or worse, two refunds. That last one compounds the refund-before-disposition window already called out: the customer is made whole early, and the ERP is the only place the double shows up.

Test it by posting RMA-4471902 through the ERP interface and then asserting on the *posted* state, not the acknowledgement — poll the ERP for the inventory record and the refund record against `OL-88231-3` and fail if either is still in staged status past a defined SLA, so "accepted but stuck" is a test failure rather than a silent gap. Assert field-by-field on what the ERP actually stored: listing grade must reflect `B`, not a defaulted `NEW`, and the stock location must be the open-box bin implied by `RESELL_OPEN_BOX`; run the same assertion against a `condition_grade` the ERP interface doesn't map cleanly and confirm it errors rather than defaulting. For duplicates, deliver the RMA-4471902 payload twice with the same `rma_id` + `order_line_id` and assert exactly one inventory unit and one refund exist afterward, with the second delivery either deduplicated or surfaced as a rejection you can see; then repeat with the second copy arriving while the first is still staged, since that's the ordering most likely to slip past a naive dedupe. Finally, reconcile the disposition store against the ERP for RMA-4471902 as a standing check — if the engine says `RESELL_OPEN_BOX` and the ERP holds no posted inventory record, that mismatch should be an alert, not something found at quarter-end.

CP1 — Capture & Transform

Returns Intake (POS/Web) → Condition/Barcode Scan

CP2 — Event Routing & Execution

Disposition Scoring Engine → Resell Channel Adapter

CP3 — Target Delivery

ERP Inventory & Refund Record

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.