← All posts

Aug 15, 2026

Event-Driven SaaS → ERP Integration

How it works

The tempting places to look are the two ends: does the RFID SaaS platform emit an event when a tag moves, and does the ERP journal store eventually show the transfer. Both are cheap to verify and both are misleading, because neither端 tells you what the pipeline did to the data in between. The real complexity begins the moment the webhook receiver accepts an event but does not carry the payload — a separate payload fetch function has to go back and retrieve it. That split is the first place where the flow stops being one thing and becomes two independently failing things: an event can be received and acknowledged while its payload is never fetched, and nothing at either endpoint will look wrong. From there the data lands in a landing zone and is picked up by a batch transform function, which means the unit of work changes shape — individual events become a batch, and the correspondence between what the SaaS platform emitted and what moves forward is no longer one-to-one. A QA engineer who only asserts "event in, journal out" has no way to detect an event that was dropped between the receiver and the fetch, or one that was absorbed silently into a batch that transformed incorrectly.

The second concentration of complexity is the routing and adaptation layer, and it deserves attention before the ERP does. The router topic publishes onward to a store transfer queue, which means delivery is decoupled and ordering, duplication, and retry behavior are now properties of the infrastructure rather than of the business logic. The store transfer adapter then has to turn a transformed batch record into something the ERP OData API will accept — this is where schema mismatch, field mapping, and rejection semantics actually live, and where a partial batch failure either surfaces or disappears. Even a successful OData call does not mean the work is done: the record enters the ERP journal store, but it is the ERP posting batch job that finalizes it, on its own schedule. That final asynchrony is why endpoint-first testing produces false negatives — a tester checking too early sees nothing posted and blames the adapter, when the adapter succeeded and the batch job simply had not run. Test the receiver-to-fetch handoff, the batch boundary, the queue's redelivery behavior, and the adapter's rejection handling first; the endpoints will confirm what those four already told you.

Caveats — what breaks in practice

The failure mode most likely to blindside a team is the one where the fetch fails after the notification has already been acknowledged, silently losing the event. Every other item on this list leaves evidence somewhere: rate limits and token expiry throw errors, schema drift breaks mapping, poison messages halt processing loudly, dead-lettered messages at least land in a DLQ even if nobody alerts on it. The acknowledged-then-failed fetch leaves nothing. The vendor has marked delivery successful and does not retry; the receiver has no payload to persist, no message to dead-letter, no record to reconcile against. The event does not appear as a failure count, a stuck staged record, or a duplicate — it appears as an absence, and absences are exactly what monitoring built around processed volume cannot see.

The assumption that makes it surprising is that acknowledgment transfers custody of the data. Teams reasonably read a 2xx on the webhook as "we have it now," because in nearly every other hop in this architecture that inference holds — a message on the queue is durable, a record accepted by the target exists somewhere, even if it is stuck staged forever or transformed by target-side defaults. Here the ack only confirms receipt of a notification, not of a payload, and the payload arrives on a separate fetch that can fail independently against rate limits, expired tokens, or pagination. Once the vendor's no-retry policy is layered on top, the acknowledgment is a promise the receiver cannot keep. For testing, this means webhook-handling tests that assert on response codes are asserting the wrong thing; the case worth building is an ack followed by an injected fetch failure, and the assertion has to be that the event is recoverable or at minimum visibly counted — not that the endpoint returned 200. It also means the only trustworthy detector in production is a reconciliation count against the source, since no error path will raise its hand.

How to test this end to end

The Store Transfer path carries a risk score of 36, and that number is not an abstraction — it is the sum of two failure modes that are invisible to any test that walks this integration front-to-back and calls each checkpoint done before moving on. The Payload Fetch Function can fail after the notification has already been acknowledged, which means the event is silently lost: nothing errors, nothing queues, nothing alerts. The ERP Posting Batch Job can stage a record perfectly and still miss the SLA on applying the business effect, which means the write looks successful and the inventory is wrong. Both failures produce green signals at the checkpoint where they occur. That is precisely why sequential checkpoint testing fails here — it is optimized for finding failures that announce themselves, and neither of these does. Testing the Store Transfer path end to end, first, is the only ordering that puts the two silent failures under observation before the noisy, self-reporting parts of the system consume the schedule.

Start at CP1, capture and transform, but start with the intent of proving loss rather than proving delivery. Correct here means the RFID SaaS Platform emits source events for every business action, with stable IDs and complete payloads retrievable via the vendor API. Validate it by triggering a known business action in the vendor UI and confirming the notification arrives — that is the cheap half. The half that matters on this path is comparing vendor-side event log counts against received counts, because that reconciliation is the only mechanism described here that can detect an event the Payload Fetch Function dropped after acknowledging the webhook. A single triggered action will pass whether or not fetch is reliable. Count reconciliation across a run is what turns a silent loss into a visible discrepancy, and it must be run against the Store Transfer path specifically, not as a general system-wide sanity pass, because that is where the risk is concentrated.

CP2, event routing and execution, is the one checkpoint on this path that behaves honestly, and testing it early pays off by removing it as an explanation for later anomalies. Correct means the Store Transfer Queue consumes messages in order, exactly once from the consumer's perspective, with the DLQ empty. Validate by checking queue depth and DLQ depth before and after a test run, and by stopping the Store Transfer Adapter, enqueuing messages, restarting, and verifying all process and none expire. The value of doing this before CP3 and CP4 is diagnostic leverage: if the queue and DLQ are clean and messages survive a consumer restart, then any missing record downstream did not vanish here, and attention moves correctly upstream to fetch or downstream to posting. Run this out of order and every later gap becomes a three-way argument.

CP3, target delivery, is where the temptation to declare victory is strongest and where the eventual-consistency gap begins. Correct means the ERP OData API returns 2xx and the created or updated record is immediately visible via a read-back query. Validate by reading the record back through the same API immediately after the write, and by submitting boundary values — zero quantity, max length strings — to confirm acceptance and rejection behave correctly. Note what this checkpoint proves and what it does not: a successful read-back proves the record is staged, not that the business effect exists. A test plan that stops at CP3 because the API returned 2xx and the record came back has confirmed exactly the condition the ERP Posting Batch Job risk describes — record staged, effect delayed past SLA. CP3 passing is the setup for the highest-risk failure on this path, not evidence against it.

CP4, target posting and application, is why the whole path has to be walked as one unit. Correct means the ERP Posting Batch Job picks up staged records on schedule and applies the business effect, such as updating on-hand inventory. Validate by checking the business effect — the on-hand quantity — after the next job run, not merely that a record was created, and by reading the job execution log for failures and for durations approaching the timeout limit. Durations near the timeout are the leading indicator of the SLA gap: the job that finishes just inside the limit today is the job that misses it under load, and the log is where that is legible before it becomes an incident. Because this checkpoint is gated on a schedule, it also sets the tempo of the whole test — you cannot validate it in the same sitting as CP3, which is a further argument for beginning with this path rather than reaching it last with time exhausted.

The ordering argument, then, is not that the other checkpoints do not matter. It is that on the Store Transfer path the two highest-risk components sit at opposite ends — Payload Fetch Function at the entry, ERP Posting Batch Job at the exit — and both fail without complaint. Only a full traversal of this path, with count reconciliation at the front and business-effect verification plus job-duration inspection at the back, closes both. Test it first, prove events are not silently lost and inventory actually moves within SLA, and everything else can be worked through afterward against a known-good spine.

CP1 — Capture & Transform

RFID SaaS Platform → Webhook Receiver → Payload Fetch Function → Landing Zone → Batch Transform Function → Router Topic

CP2 — Event Routing & Execution

Store Transfer Queue → Store Transfer Adapter

CP3 — Target Delivery

ERP OData API → ERP Journal Store

CP4 — Target Posting & Application

ERP Posting Batch Job

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.