Aug 27, 2026
Saga-Based Distributed Transaction
How it works
The obvious reading of this chain is that it is a pipeline: the client calls the gateway, the gateway hands off to the orchestrator, and payment, inventory, and shipping fire in order until shipping succeeds. That reading is wrong in the way that matters. Service A, B, and C are not stages of one transaction; they are three independent local commits, each of which is already durable the moment it returns. The orchestrator is not a router between them — it is the only thing in the flow that knows a business transaction exists at all, and the Transaction State Store is the only place that knowledge is written down. Everything the flow can promise about correctness therefore lives in two components that the client never touches and that neither the gateway nor the individual services can reconstruct. When inventory fails after payment has committed, no amount of correctness inside Service A saves you; what saves you is the orchestrator reading its state from the store, deciding that payment must be compensated, and driving that reversal to completion. The forward path is the cheap part. The decision layer is the product.
So a QA engineer who starts at the endpoints — gateway contract tests, per-service happy paths, a final assertion that shipping was created — has tested the least interesting third of the system. The tests that earn their keep target the orchestrator's behavior against the state store: does a failure at C leave a state record that correctly drives compensation back through B and A, and does the orchestrator resume from that record rather than from nothing if it dies mid-saga? Because the state store is what survives a restart, a stale, missing, or half-written state entry converts a recoverable failure into a permanently inconsistent order, with payment taken and no shipment. The Event/Audit Log is the second place to look, and for a different reason: it is the only artifact that shows the sequence the orchestrator actually chose, so it is where you verify that compensations were attempted in the right order and not silently skipped. Test the orchestrator's failure branches and the durability of what it writes first; the endpoints will still be there when you get to them.
Caveats — what breaks in practice
The failure mode most likely to blindside a team is the success response returned for a record that was never actually persisted. Every other item on the list at least announces itself somewhere: duplicates on retry show up as double rows, poison messages halt processing, timeouts on oversized batches throw, oversized payloads are rejected. Those produce a signal, even if the signal is ugly or late. A false success produces nothing. The saga advances, the next step commits for real, and the orchestration reaches a "completed" state built on a step that didn't happen. It is indistinguishable from correctness until something downstream reads the missing record — and in a system that already tolerates replication lag between primary and replica reads, the first person to notice will reasonably assume they hit stale data and wait for it to resolve.
The assumption that makes it surprising is that the acknowledgement is the commit — that a saga step's success response is a durable fact the compensation logic can trust. Sagas are built entirely on that trust: there is no distributed lock, no two-phase commit, only a chain of local transactions and compensating actions triggered by reported failure. A step that lies about succeeding never reports failure, so compensation never fires, and the saga has no mechanism to detect the gap on its own. This is the same assumption that makes silent gaps — changes saved without emitting an integration event — dangerous, and the two compound: one loses the write, the other loses the notice of it. For testing, this means acknowledgement-level assertions are worthless as saga step verification. Every step that claims success has to be confirmed by independently reading the record back from the authoritative store, not the replica, and compensation paths need to be exercised against a step that returns success and persists nothing — a case that will not appear unless you deliberately construct it.
How to test this end to end
A customer taps "Place Order" in the mobile app, which POSTs to the API Gateway with `orderId: "ORD-2024-88431"`, `customerId: "CUS-55210"`, `paymentMethodToken: "tok_visa_4242"`, `currency: "USD"`, `orderTotal: 249.97`, and a `lineItems` array of three entries (`SKU-CHAIR-9`, qty 1; `SKU-CUSHION-2`, qty 2; `SKU-GIFTWRAP-0`, qty 1 at price 0.00), plus an `Idempotency-Key: 9f3c-ORD-2024-88431` header. The Gateway authenticates the device credential, validates the payload, and hands a normalized saga-start command to the Saga Orchestrator, which is the thing that will subsequently drive Payment, then Inventory, then Shipping in sequence — and, if Shipping fails, drive the compensations back through Inventory and Payment. Everything downstream of this checkpoint depends on ORD-2024-88431 arriving at the orchestrator exactly once, complete, and in a shape the orchestrator understands.
What could go wrong here is mostly invisible at the moment it happens. The app can retry on a transient gateway timeout and start two sagas for ORD-2024-88431, meaning `tok_visa_4242` gets charged twice with two separate compensation obligations. Or the Gateway accepts and acknowledges the order but never emits the saga-start command — the customer sees a confirmation and no service ever runs. Edge-case mapping is a real risk with this record: `SKU-GIFTWRAP-0` at price 0.00 or a null `giftMessage` can fail transformation and become a poison message that stalls the queue, and if the orchestrator's command schema drifts (say `orderTotal` becomes `amount.value`) the saga starts with a missing total and Payment gets a garbage request. Under burst load, throttling can drop ORD-2024-88431 silently, and an expired device credential can be rejected with no alert. Test it by submitting ORD-2024-88431 twice with the same `Idempotency-Key` and asserting exactly one saga instance exists at the orchestrator; assert that every accepted order at the Gateway produces a corresponding saga-start command (reconcile counts, not just spot checks); replay ORD-2024-88431 with `SKU-GIFTWRAP-0` priced 0.00 and null optional fields and confirm it maps cleanly rather than parking the queue; fire a burst of a few hundred copies of the order and confirm throttled requests surface as errors to the client rather than vanishing; and send it with an expired device credential and verify the rejection raises an alert, not just a 401.
Once the saga-start command for ORD-2024-88431 lands, the orchestrator drives the sequence and writes a state row at every hop. Payment is called first with the `paymentMethodToken: "tok_visa_4242"` and `orderTotal: 249.97` in `USD`, and returns `paymentId: "PAY-77301"`, `status: "AUTHORIZED"`; the orchestrator writes `{sagaId: "SAGA-ORD-2024-88431", step: "PAYMENT", state: "COMPLETED", compensationRef: "PAY-77301"}` to the Transaction State Store. Inventory is next, reserving `SKU-CHAIR-9` qty 1, `SKU-CUSHION-2` qty 2, and `SKU-GIFTWRAP-0` qty 1, returning `reservationId: "RES-41988"`, which is likewise recorded as `step: "INVENTORY", state: "COMPLETED"`. Shipping is called last with the destination and reservation, and if it returns a failure — say `SHIP_UNSERVICEABLE_ZIP` — the orchestrator reads back the state store and drives compensations in reverse: release `RES-41988`, then void or refund `PAY-77301`. Every one of those compensating calls depends on the state store having actually persisted `PAY-77301` and `RES-41988` at the moment they succeeded, because the state store is the only record that those steps ever happened.
The dangerous failures here are the quiet ones. Payment can return `AUTHORIZED` while the state write for `PAY-77301` never actually commits despite a success response, so when Shipping fails there is no `compensationRef` to refund and the customer is charged 249.97 with nothing shipped. Inventory can choke on `SKU-GIFTWRAP-0` at price 0.00 during validation and fail mid-reservation, leaving `SKU-CHAIR-9` reserved but no `RES-41988` returned. If the orchestrator writes state to a primary but reads compensation refs from a replica, replication lag can make `RES-41988` invisible during a fast-failing Shipping call and skip the release entirely. A migration that renames `compensationRef` to `compensation_ref` breaks the compensation reader silently — the saga looks complete while nothing is undone. And a long-running Shipping call holding the saga row can create lock contention that times out concurrent sagas, while connection pool exhaustion under batch load surfaces as timeouts in Payment rather than where the pressure actually is. Test it by forcing Shipping to return `SHIP_UNSERVICEABLE_ZIP` for ORD-2024-88431 and asserting that both `RES-41988` is released and `PAY-77301` is refunded, not just that the saga is marked failed; kill the state-store write immediately after Payment returns `AUTHORIZED` and confirm the orchestrator does not silently proceed without a `compensationRef`; run ORD-2024-88431 with reads pinned to a replica under induced lag and verify compensations still find `RES-41988`; replay the saga against a post-migration schema and assert the compensation reader fails loudly rather than reporting success; and drive a few hundred concurrent copies of the order to confirm pool exhaustion produces explicit saga failures with intact compensation, not orphaned authorizations.
Downstream of the orchestrator's state writes, every hop also emits an event to the audit log sink, and for ORD-2024-88431 that stream is the only durable narrative of what the saga actually did. A completed-then-compensated run lands roughly five records: `{sagaId: "SAGA-ORD-2024-88431", step: "PAYMENT", state: "COMPLETED", compensationRef: "PAY-77301", amount: 249.97, currency: "USD", emittedAt: "2024-06-11T14:22:03.114Z"}`, then the same shape for `step: "INVENTORY"` with `compensationRef: "RES-41988"`, then `step: "SHIPPING", state: "FAILED", errorCode: "SHIP_UNSERVICEABLE_ZIP"`, followed by the two compensation events `step: "INVENTORY", state: "COMPENSATED"` and `step: "PAYMENT", state: "COMPENSATED"`. This is what an ops engineer reads at 2am to answer "was the customer's 249.97 actually refunded," and what finance aggregates when reconciling authorized versus refunded totals for the day.
The failure modes bite precisely because the sink is write-and-forget. If the sink's schema was defined before compensation existed and has no `compensationRef` column, the field is silently dropped on ingest — the audit rows for SAGA-ORD-2024-88431 still say PAYMENT/COMPLETED and PAYMENT/COMPENSATED, but `PAY-77301` is nowhere in the log, so nobody can tie the refund to the authorization; same hazard if `errorCode` is dropped and `SHIP_UNSERVICEABLE_ZIP` vanishes, leaving an unexplained failure. The other risk is replay: if the saga is re-driven or the emitter retries, a second `PAYMENT/COMPLETED` event with `amount: 249.97` lands and daily authorized-total aggregation double-counts to 499.94 against a customer who was charged once and refunded once. To test, run the forced-`SHIP_UNSERVICEABLE_ZIP` scenario for ORD-2024-88431 and assert on the sink side — not the orchestrator side — that all five events are present and that `compensationRef` reads back as exactly `PAY-77301` and `RES-41988` rather than null; add a schema-drift case where the sink is missing `compensationRef` and confirm ingest rejects loudly instead of accepting a truncated row. Then replay the same saga twice with identical `sagaId` and assert the sink holds one event per `(sagaId, step, state)` and that the aggregated authorized total for SAGA-ORD-2024-88431 stays at 249.97, not 499.94.
CP1 — Capture & Transform
Client / Channel App → API Gateway → Saga Orchestrator
CP2 — Flow Stage
Service A (Payment) → Service B (Inventory) → Service C (Shipping) → Transaction State Store
CP3 — Sink Delivery
Event / Audit Log
Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.