← All posts

Sep 5, 2026

Cart Abandonment Trigger

How it works

The obvious reading of this flow is that the Cart/Session Service knows when a cart is abandoned and the Marketing Automation Platform sends the email, with a queue in the middle doing nothing more interesting than waiting. That reading is wrong, and it misdirects testing effort to the two endpoints that are actually the easiest parts to verify. A cart service can be checked against its own state; a marketing platform can be checked against whether a message went out. Neither of those endpoints holds the decision that defines the feature. The real question — is this cart abandoned? — is not answered by the cart service, which only knows that nothing has happened yet, and it is not answered by the marketing platform, which receives a verdict rather than reaching one. It is answered by the Abandonment Timer Queue, and specifically by the interval between the session's last activity and the moment the queue releases the trigger. That gap is where the product's meaning lives, and it is the only component in the chain whose correct behavior is defined by the passage of time rather than by a value you can inspect.

That makes the timer queue the place a QA engineer should start, because time-dependent correctness fails in ways endpoint assertions never surface. The interesting cases are all about what happens to a pending timer when the cart underneath it changes: a customer who adds an item after the timer is armed, a customer who completes the purchase before it fires, a session that resumes and then goes quiet again. In each case the cart service's state is correct and the marketing platform behaves correctly given what it was told — and the outcome is still wrong, because a timer fired against a cart that no longer justified it. Testing the endpoints will pass all of these. So will testing the happy path, which is precisely the scenario where nothing changes during the wait. Attention belongs on whether the queued trigger is cancelled, reset, or re-evaluated when the underlying cart moves, and on what the marketing platform receives when a trigger fires against stale state, because that is the only interface in this flow where a component acts on information it captured earlier rather than information it holds now.

Caveats — what breaks in practice

The failure mode most likely to blindside a team is duplicate cart-updated events restarting the timer repeatedly, delaying the message indefinitely. Every other item on this list eventually announces itself. A duplicate send after a retry produces a customer-visible second message. A timer that fires after purchase produces an angry "I already bought this" ticket. Schema drift breaks a consumer loudly. Validation rejections and throttling leave error records behind. Even silent gaps and phantom writes eventually show up as a record someone goes looking for and cannot find. But a timer that keeps resetting produces no error, no rejected record, no duplicate, and no wrong message — it produces nothing at all, on a path where nothing is also what a normal completed purchase looks like. The abandonment funnel just runs a little cooler than expected, and that gets attributed to the market.

The assumption that makes it surprising is that duplicate events are a solved and benign problem — that because retries are expected and the system tolerates them, receiving the same cart-updated event twice is equivalent to receiving it once. That holds for state writes; it does not hold here, because the timer treats each arriving event as evidence of fresh shopper activity rather than as a restatement of a known fact. Deduplication logic that protects the outbound side, which is where teams put it after being burned by duplicate messages on retry, does nothing for the inbound side where the damage actually occurs. For testing, this means duplicate-delivery scenarios cannot be verified only by asserting "exactly one message sent"; the case that matters is replaying the same cart-updated event across the threshold window and asserting the message still fires on the original schedule. And it means the suite needs a negative-time assertion — that a message is emitted by a deadline — because a test that only checks correctness of messages that were sent will pass perfectly against a trigger that has silently stopped sending anything.

How to test this end to end

Picture shopper `cust_88431` adding a second pair of running shoes at 14:02:11 UTC. The Cart/Session Service writes the line item to its own store and then emits a `cart.updated` integration event onto the Abandonment Timer Queue carrying something like `event_id: "evt_7f3a1c"`, `cart_id: "cart_2291043"`, `customer_id: "cust_88431"`, `session_id: "sess_aa19"`, `item_count: 2`, `cart_total: 218.00`, `currency: "USD"`, `updated_at: "2024-03-14T14:02:11Z"`, `schema_version: "1.3"`. That message is the entire basis for the delay downstream: the timer that will decide, hours later, whether `cust_88431` gets a marketing message exists only because this event was published, with this timestamp, in this shape.

What goes wrong here is mostly invisible on a happy-path demo. If the service commits the line item but the publish fails, `cart_2291043` never gets a timer at all — a silent gap nobody notices because nothing fired. If the publish is retried after a transient broker failure, two `cart.updated` events land for the same `event_id: "evt_7f3a1c"`, and since each cart-updated event restarts the timer, duplicates can push `cust_88431`'s message out indefinitely — or, worse, the same duplication pattern means a cancel-side event for a completed purchase has to compete with an already-scheduled timer. If the service's own model changes and `cart_total` becomes a nested `{amount, currency}` object without bumping `schema_version` past `1.3`, the consumer reading the queue drifts off the contract. And if `updated_at` is emitted in local time rather than UTC, the threshold window for `cust_88431` shifts by hours in either direction. To test: assert that every committed cart mutation for `cart_2291043` produces exactly one queue message, by forcing a broker failure mid-publish and confirming the retry is deduplicated on `event_id: "evt_7f3a1c"` rather than enqueued twice; fire ten rapid updates to `cart_2291043` and confirm the timer restart behavior matches the documented intent rather than starving the message forever; contract-test the emitted payload against `schema_version: "1.3"` so a model change fails the build instead of the consumer; and assert `updated_at` is always UTC by running the producer under a non-UTC system clock and checking the emitted value still reads `2024-03-14T14:02:11Z`.

Hours later the timer for `cart_2291043` expires without a purchase event, and the handoff to the Marketing Automation Platform is where the abandonment decision finally becomes a real outbound record. The delivery payload carries the shopper identity and cart facts forward from `evt_7f3a1c` — something like `contact_key: "cust_88431"`, `campaign: "cart_abandon_24h"`, `cart_id: "cart_2291043"`, `cart_total: 218.00`, `currency: "USD"`, `item_count: 2`, `cart_updated_at: "2024-03-14T14:02:11Z"` — and the platform validates it, accepts it into its contact/journey store, and returns a success acknowledgment. That acknowledgment is the only signal our side gets that `cust_88431` is actually enrolled and will be messaged.

The trouble is that acknowledgment is not proof. The platform's validation may reject edge-case values outright — an empty `contact_key`, a `cart_total: 0.00` left over from a cart emptied to a single removed item, an `item_count` that arrived as a string, or a currency the campaign isn't configured for — and if that rejection is logged but not surfaced, `cust_88431` silently never gets the message, exactly the same observable outcome as the publish having failed back at CP1. Under batch load the platform throttles, so a burst of expiring timers gets partially shed and `cart_2291043` is one of the dropped ones while the overall job still reports success. And the nastiest case is a 200-style success response where the record never persists: the journey entry for `cust_88431` simply isn't there when you look, so no message fires and nothing anywhere says it failed. To test, deliver the `cart_2291043` record with deliberately hostile field values — `cart_total: 0.00`, `item_count: 0`, a blank `contact_key`, a `currency: "XYZ"` — and assert each rejection is raised as a visible failure rather than swallowed; push a batch of several thousand expiring timers containing `cart_2291043` and assert that every single one is either enrolled or explicitly reported as throttled and retried, with none unaccounted for; and after each accepted delivery, read back the record from the platform by `contact_key: "cust_88431"` and `cart_id: "cart_2291043"` and fail the test if the enrollment isn't actually queryable, so a success response with no persisted record is caught rather than trusted.

CP1 — Capture & Transform

Cart/Session Service → Abandonment Timer Queue

CP2 — Target Delivery

Marketing Automation Platform

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.