Aug 26, 2026
Master Data Management Hub-and-Spoke
How it works
The obvious places to test in this flow are the ends: does Source System A's product data land in staging, and does the ERP end up with what the hub published? Both are cheap to verify and both are nearly useless as a measure of correctness, because neither can tell you whether the golden record is right — only whether something arrived. The real complexity sits in the Match & Merge Engine and the Data Steward Review Queue, and specifically in the seam between them. The engine makes probabilistic decisions about which incoming product records refer to the same real thing, and every one of those decisions either creates a golden record, collapses two distinct products into one, or splits one product across two. The review queue exists precisely because the engine is not trusted to be right, which is an admission built into the architecture: the system's designers knew match and merge would produce outcomes requiring human adjudication. A QA engineer who accepts that admission should conclude that this is where defects live, not at ingestion.
That conclusion sharpens once you follow what happens after a steward acts. The Golden Record Store feeds the review queue, and steward decisions feed back into what the hub considers authoritative — which then goes out over the Publish Event Bus to ERP. So a bad merge is not a contained error; it is an error that gets ratified, published as an event, and synced downstream into a system that will treat it as the truth about a product. The endpoints are where you observe the damage, but they are not where you can prevent it, because by the time the event bus has fired, the ERP has no basis to question what it received. Testing effort belongs on merge outcomes under ambiguous source data, on whether records that should reach a steward actually reach the queue rather than auto-merging past it, and on whether a steward's correction actually propagates as a published event rather than stalling in the hub. Verifying that ERP matches the hub proves only that sync works; it says nothing about whether the hub was ever correct.
Caveats — what breaks in practice
The failure most likely to blindside a team is the silently black-holed event: a new event type published into the hub that matches no subscription filter, or an existing type quietly discarded by a filter mismatch. Nothing errors. There is no dead-letter message to inspect, no retry storm, no timeout, no poison message halting a queue — all of which announce themselves. The publisher gets a successful write to the hub, the spoke that never needed the event behaves normally, and the one spoke that did need it simply never sees a record that, as far as every dashboard is concerned, exists. Compare this to DLQ accumulation without alerting, which is also quiet but at least leaves a physical pile of evidence to find; a filter mismatch leaves nothing at all. The gap surfaces days later as a data discrepancy someone stumbles into, not as an incident.
The assumption doing the damage is that publishing to the master hub is the same thing as distributing to the spokes — that once the golden record is written, propagation is a solved, automatic property of the architecture. It is reasonable to believe this, because that is exactly what hub-and-spoke is sold as. But subscription filters are configuration living outside the publisher's change, so adding an event type or an event subtype is a two-sided change that only looks one-sided. For testing, this means contract tests on the publish side prove nothing; the assertion has to be that every published type is claimed by at least one subscription, and that unmatched publishes are counted and alerted rather than dropped. Test the routing table as a first-class artifact, and add negative coverage for the new-type case specifically, because field mapping errors for specific event subtypes hide in exactly the same blind spot — the payload arrives, the shape is wrong for that subtype, and nobody is watching the seam between publish and subscribe.
How to test this end to end
Picture a nightly delta from Source System A, the PIM that owns product data: a webhook fires with `sku: "PRD-88431"`, `product_name: "Acme 12V Cordless Drill"`, `uom: "EA"`, `list_price: 129.00`, `status: "ACTIVE"`, `source_updated_at: "2024-03-11T02:14:07Z"`. The Ingestion & Staging Layer fetches the full payload from the vendor API, lands it in staging keyed by source system plus SKU, and the Match & Merge Engine compares it against two other candidate records already in the hub — one with `sku: "PRD-88431"` and one legacy record `sku: "88431-ACME"` carrying the same GTIN. Survivorship rules score the pair at, say, 0.78 — above noise, below the auto-merge threshold — so instead of writing a merged golden record straight through, the engine writes the candidate cluster to the Golden Record Store and drops a task into the Data Steward Review Queue. Only after a steward confirms the merge does the golden record for `PRD-88431` hit the Publish Event Bus and fan out to every subscribing downstream system.
The dangerous version is quieter than an outright crash. Suppose the vendor's March release adds `uom_code` and starts sending `uom: null`, or `status: "ACTIVE_LIMITED"` — an enum the mapping layer has never seen. A mapping error on that edge-case value can push `PRD-88431` through with an empty unit of measure, and because the record now looks *less* distinct from `88431-ACME`, the match score can drift above threshold and auto-merge without ever reaching the steward queue — a merge error that publishes. Or the reverse: the auth token expires mid-fetch, the batch is retried, and `PRD-88431` is redelivered after a consumer timeout, creating a duplicate cluster; or the oversized nightly batch is split and the drill's line item lands in neither half. Worst case, the new `ACTIVE_LIMITED` event type matches no subscription filter downstream and is silently black-holed, so nobody notices until the hub and the spokes disagree — and by then the fix is a correction in the hub plus a re-sync everywhere, not a patch in one system.
To test it, replay `PRD-88431` through staging with a deliberately drifted payload — `uom: null` and `status: "ACTIVE_LIMITED"` — and assert three things: the ingestion layer either rejects or quarantines the unmapped enum rather than defaulting it, the match score against `88431-ACME` still lands the pair in the Data Steward Review Queue rather than auto-merging, and nothing appears on the Publish Event Bus until a steward decision is recorded. Then re-send the identical `PRD-88431` webhook twice with the same `source_updated_at` to confirm the second delivery is deduplicated instead of spawning a second candidate cluster, and kill the consumer mid-batch to check that the drill's line item is present exactly once after recovery — not lost to a batch split, not duplicated by redelivery. Finally, publish the confirmed golden record and verify every subscriber actually received `PRD-88431`; if a subscription filter drops it, the test should fail loudly rather than leaving the record on a DLQ nobody is watching.
Once the steward confirms the merge, the golden record for `PRD-88431` lands on the Publish Event Bus and the ERP subscriber picks it up — say an envelope carrying `event_type: "golden_record.updated"`, `mdm_id: "GR-40217"`, the surviving `sku: "PRD-88431"`, `product_name: "Acme 12V Cordless Drill"`, `uom: "EA"`, `list_price: 129.00`, `status: "ACTIVE"` — and translates it into the ERP's material master call, mapping `uom` onto the ERP's base unit field and `list_price` onto its standard price. That mapping is per-subtype: a `golden_record.merged` event also has to tell the ERP that legacy material `88431-ACME` is now an alias of `GR-40217`, and if the merge subtype maps only the surviving SKU's fields, the ERP keeps both materials live and quietly diverges from the hub. Meanwhile the ERP session token expires mid-call and the write returns 401; the connector retries, the ERP stays down for twenty minutes, and the retry loop hammers it into a storm. Worse, if a second event for `GR-40217` — a later price change to `139.00` — is processed concurrently with the redelivered merge event, the two writes race and the ERP can settle on the older `129.00` while the hub says `139.00`. None of that raises an alarm in the hub; the golden record was published successfully, so from the hub's point of view CP2 succeeded.
To test it, drive the confirmed `PRD-88431` merge onto the bus and assert the ERP connector emits both the material update for `GR-40217` and the alias/deactivation call for `88431-ACME`, then repeat with a `golden_record.updated`-only event to prove the two subtypes take different mapping paths rather than sharing one lazy handler. Expire or revoke the ERP session token immediately before delivery and confirm the connector re-authenticates and completes the write once, instead of consuming the message and dropping it; then hold the ERP in a hard-failure state and check that retries back off and land `PRD-88431` on a monitored DLQ with an alert, not an unbounded loop. For the race, publish the merge event and a follow-up price change to `list_price: 139.00` for `GR-40217` within the same second and assert the ERP ends at `139.00` every run — ordering or optimistic-concurrency enforced, last-write-wins by event timestamp, not by whichever thread finished first — and finally reconcile the ERP material record against the hub's golden `PRD-88431` field by field, so a silent mapping gap fails the test rather than surfacing weeks later as a pricing dispute.
CP1 — Projection & Apply
Source System A (product data) → Ingestion & Staging Layer → Match & Merge Engine → Golden Record Store (MDM Hub) → Data Steward Review Queue → Publish Event Bus
CP2 — Event Routing & Execution
Downstream Sync — ERP
Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.