Sep 7, 2026
Offline-First POS Sync
How it works
The obvious places to look in this flow are the two endpoints: the POS terminal where a cashier rings up a sale, and the central order and inventory system where the number finally lands. Both are easy to test because both are deterministic — a tap produces a line item, a posted order produces a stock decrement. But neither endpoint is where this architecture earns its name. "Offline-first" means the terminal is designed to succeed without the central system existing at all, and the local transaction cache is designed to accept writes that no authority has yet validated. The moment you accept unvalidated writes, you have created a second source of truth, and the sync/conflict-resolution engine is the only component in the flow whose job is to reconcile two truths into one. Every hard problem in this pipeline is that reconciliation problem wearing a different costume: the same physical item sold from two terminals whose caches never saw each other, an order cached before a price change and synced after it, a cache replaying a transaction the central system already recorded.
So a QA engineer's first attention belongs on the conflict-resolution engine, and specifically on the cases the endpoints cannot generate on their own. Testing the terminal tells you a sale is captured; testing the central system tells you a posted order updates inventory; neither tells you what happens when two caches arrive holding incompatible claims about the same stock. Those scenarios only exist as inputs to the engine, which means they have to be constructed deliberately — divergent caches, out-of-order arrival, replayed syncs — rather than discovered by exercising the happy path from the terminal forward. The endpoints will look correct in almost every test you write against them, because they are correct in isolation; the defect surface lives in the space between them, and it is only observable in what the central inventory reads after the engine has decided which of two competing local truths survives.
Caveats — what breaks in practice
The failure most likely to blindside a team is the crash mid-sync that leaves some transactions applied and others lost, with no record of which. Every other item on this list announces itself: duplicate events on retry show up as inflated counts, out-of-order replay corrupts running totals visibly, a record stuck in staged state forever is queryable, schema drift eventually throws or produces obviously wrong fields. A partial sync with no record of what landed produces a result set that is internally consistent and simply incomplete — the register reports success for what it sent, the backend reports success for what it received, and the gap between those two numbers exists nowhere. It resembles the silent-gap failure, changes saved without emitting an integration event, except that it is created by the recovery path rather than the write path, which means the safeguards a team builds for silent gaps don't catch it.
The assumption that makes it surprising is that the sync batch is atomic or, failing that, resumable — that after a crash you can either roll back and resend or ask the target what it already has. That expectation is entirely reasonable; it's how retry handling for duplicate events is justified in the first place, since idempotent replay is safe only if you know what to replay. Offline-first removes the arbiter: the register's queue and the backend's applied set are two independent records with no shared transaction boundary, and the crash destroys the only thing correlating them. For testing, this means the interesting case is not the clean disconnect but the interrupted reconnect, and the assertion is not "did the sync succeed" but "can the system enumerate exactly which transactions were applied." Kill the process partway through a multi-transaction sync, reconnect, and demand that enumeration — a system that can't produce it will hide the same failure in production behind a green success path, and the inventory conflict between two registers over a low-stock item will surface later as the symptom rather than the cause.
How to test this end to end
Picture register REG-04 at the Northgate store at 14:22 on a Saturday, uplink down. A customer buys the last SKU-88213 "Hydro Flask 32oz Slate" — the terminal writes a local sale row `txn_id: REG04-20240615-000871`, `sku: SKU-88213`, `qty: 1`, `unit_price: 44.95`, `local_seq: 871`, `captured_at: 2024-06-15T14:22:07-07:00`, `sync_state: PENDING`, and decrements its own cached `on_hand` for that SKU from 1 to 0. The same write is supposed to append an outbound integration event into the sync queue — `SaleRecorded` v3, carrying `txn_id`, `sku`, `qty`, `local_seq`, and the register's `device_id: REG-04` — so that when the link returns the conflict-resolution engine has something to replay. Two aisles over, REG-07 is also offline and sells the same physical-looking last unit as `REG07-20240615-000412`, cached `on_hand` also going 1 to 0. Both registers believe they are consistent; the engine only learns otherwise at reconnect, when it must reconcile two independent decrements against a central on-hand of 1.
The dangerous variant is not the conflict itself but the silent gap: if the local cache commits REG04-20240615-000871 to the sale table but the outbound event append fails or is skipped, the register's drawer total and the customer's receipt both reflect a sale the engine will never see, so reconciliation "succeeds" while the central ledger is short one unit of revenue and one unit of shrink — and nothing errors. Closely related, a retry after a transient write can append `SaleRecorded` for 000871 twice, double-decrementing on replay, and replay that arrives out of `local_seq` order can corrupt the running drawer total. To test, drive REG-04 offline, ring 000871, then assert at the storage layer that the sale row and its outbound event exist in the same committed unit — kill the process between the two writes and confirm neither survives alone. Then reconnect both registers and assert the engine sees exactly one event per `txn_id` (force a duplicate append of 000871 and verify idempotent suppression on `txn_id`), that events apply in ascending `local_seq` per `device_id`, and that reconciling 000871 against REG07-20240615-000412 for SKU-88213 yields an explicit, logged oversell exception rather than a silently negative or silently clamped on_hand.
When the uplink comes back at 14:47, the sync engine hands `SaleRecorded` v3 for `REG04-20240615-000871` across to the Central Order & Inventory System, which accepts it into a staging table before it touches the authoritative on-hand for SKU-88213 — a staged row carrying `txn_id`, `sku`, `qty: 1`, `local_seq: 871`, `device_id: REG-04`, and a target-assigned `stage_status: STAGED`. REG-07's `REG07-20240615-000412` lands the same way moments later. Only when the central system promotes those staged rows does the real decrement happen, and that promotion is where the central on-hand of 1 meets two independent claims on the same physical unit and the oversell exception is supposed to be raised and logged. Until promotion, both rows sit there looking perfectly healthy, and the register that rang 000871 has no visibility into whether it ever moved past `STAGED`.
Three things go wrong here and none of them look like errors upstream. The staged row for 000871 can be accepted and simply never promoted — the engine reports a clean reconnect, the register clears `sync_state: PENDING`, and the sale is parked indefinitely in a state nobody polls, which is functionally the same lost unit of revenue as the missing outbound event, just moved one hop downstream. Second, target-side defaults can quietly rewrite the payload: if `unit_price` is absent from the v3 event and the central system fills a list price of 49.95, or coerces `captured_at` to server time and loses the `-07:00` offset, 000871 reconciles against the wrong day's close. Third, if the engine redelivers after a timeout, the central system can end up with two staged copies of 000871 — upstream idempotency on `txn_id` does not help if the dedupe check lives only on the register side. To test, replay 000871 and 000412 and assert both staged rows reach a terminal state, not `STAGED`, within the promotion window; deliberately stall promotion and confirm something alerts rather than the register showing a clean sync. Compare the promoted values for 000871 field-by-field against the terminal's local sale row — `unit_price: 44.95` and `captured_at: 2024-06-15T14:22:07-07:00` must survive intact, not be repopulated by defaults. Then redeliver the 000871 event twice and assert exactly one staged row exists for that `txn_id`, and that the promotion of 000871 and 000412 against a central on-hand of 1 for SKU-88213 produces one applied decrement and one logged oversell exception.
CP1 — Projection & Apply
POS Terminal → Local Transaction Cache → Sync/Conflict-Resolution Engine
CP2 — Target Delivery
Central Order & Inventory System
Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.