Sep 10, 2026
EMPI Identity Resolution
How it works
The obvious places to test this flow are its ends: did the Registration System emit an event, and did the EMPI Identity Store end up with the right golden record. Both are cheap to assert and both are nearly useless as a measure of whether identity resolution works, because neither endpoint holds a decision. The Registration System only reports what a clerk typed; the EMPI Identity Store only records what something upstream already concluded. The actual reasoning lives in the Probabilistic Match Engine and, more importantly, in the Match Candidate Store that sits between the engine's output and the Auto-Merge Adapter's action. That store exists precisely because the engine does not return yes or no — it returns candidates with scores, and something has to decide which of those scores are strong enough to act on without a human. That decision boundary is the system, and it is the only component in the chain that can be wrong in a way no endpoint assertion will reveal.
This matters because the two failure modes at that boundary are asymmetric and neither shows up as an error. If the Auto-Merge Adapter acts on a candidate that should have stayed a candidate, two distinct people collapse into one identity in the EMPI Identity Store, and the flow as drawn has no reverse path — the Registration Event Feed moves one direction only. If it declines to act on a candidate that was genuinely a match, the same person persists as two identities and nothing in the pipeline raises a complaint. A QA engineer who tests only the endpoints will see clean events going in and populated records coming out in both cases. The tests that are worth writing are the ones that pin the Match Candidate Store's contents against known-truth pairs — deliberate near-misses, shared surnames and birthdates, transposed digits — and assert not just the final merge but which candidates the Auto-Merge Adapter consumed and which it left behind. Start there, because that is where a correct-looking EMPI Identity Store is most likely to be quietly wrong.
Caveats — what breaks in practice
The failure most likely to blindside a team is the record that is accepted but sits in a staged state forever. Every other item on this list announces itself: a timeout on an oversized batch throws, a poison message halts processing, a signature validation failure rejects, a retry storm lights up the target's error rate. Staged-forever does none of that. The EMPI accepts the record, returns success, and the interface moves on clean. There is no error to alert on, no queue depth to watch climb, no failed write to retry. The only symptom is a patient identity that never resolved — which surfaces weeks later as a duplicate chart or a missing link, far from the integration point that caused it, and usually reported by a clinician rather than a monitor.
The assumption that makes it surprising is that acceptance equals resolution — that a successful response from the EMPI means the identity has been matched and committed, when in fact acceptance only means the payload was taken in for evaluation. That assumption is entirely reasonable everywhere else in this pipeline, and it is reinforced by the neighboring failure modes: partial writes on failure and processing triggered before the write completes both teach the team to distrust *failed* or *early* operations, which quietly trains them to trust the successful ones. For QE, this means acknowledgement-based assertions are worthless here. A test that posts a patient and asserts on the response is testing the wrong boundary. The check has to be a separate read against resolved state, with an explicit time bound, plus a standing production query for records that have been staged longer than that bound — because nothing in this system will ever raise its hand for them.
How to test this end to end
Picture an A04 registration from the Riverbend Clinic site, message control ID `RVB-2024-118773`, carrying local MRN `RVB-4471902` for "Jonathan R. Meier," DOB `1978-03-14`, SSN last four `8842`, address `412 Wexford Ln`. The Registration Event Feed hands that payload to the Probabilistic Match Engine, which scores it against existing identities and lands on EMPI identity `EID-00913338` ("Jon Meier," MRN `CTH-220814` from Central Hospital, same DOB, SSN match, address differing by apartment number) at a composite score of `0.94`. That score clears the auto-merge threshold, so a candidate row is written to the Match Candidate Store with status `CANDIDATE_AUTO`, and the Auto-Merge Adapter picks it up to finalize the link between `RVB-4471902` and `EID-00913338`. Note what actually crosses the checkpoint: the local MRN goes in, the resolved `EID-00913338` comes out, and those are two different identifiers that downstream consumers must not treat as interchangeable.
The dangerous thing is that the candidate row and the completed merge are separate moments. If the interface restarts mid-flight and `RVB-2024-118773` is redelivered, you can get two candidate rows for the same registration and a concurrency race where two Auto-Merge Adapter invocations update `EID-00913338` at once, leaving a partial write — the link recorded on one side but not the other. Worse, a mapping error on an edge-case value (a null middle initial, or Riverbend sending a non-standard gender enum) can drop a discriminating field from the score, pushing a genuinely different Jonathan Meier from `0.61` up past the auto-merge line and fusing two patients into one identity with no steward ever seeing it. And if the Auto-Merge Adapter's target session expires, the row sits at `CANDIDATE_AUTO` forever, looking healthy to anyone who only checks that staging happened. To test it, replay `RVB-2024-118773` twice and assert exactly one finalized link from `RVB-4471902` to `EID-00913338`, not just one candidate row; then poll the Match Candidate Store until status leaves `CANDIDATE_AUTO`, with an explicit timeout that fails the test rather than passing on staging alone. Separately, resend the same registration with the SSN and address stripped to null and confirm the score drops below the auto-merge threshold and routes to steward review instead of silently merging into `EID-00913338`.
At CP2 the thing being delivered is the finalized link itself: the Auto-Merge Adapter writes into the EMPI Identity Store so that `EID-00913338` now carries both `CTH-220814` from Central Hospital and `RVB-4471902` from Riverbend, with the demographic payload from `RVB-2024-118773` — "Jonathan R. Meier," DOB `1978-03-14`, SSN last four `8842`, address `412 Wexford Ln` — landing alongside the existing "Jon Meier" attributes on that identity. What crosses the boundary here is the local MRN and its demographics going in and the resolved `EID-00913338` being what the store now answers with; the store is the authority that says those two MRNs are one person, and any downstream consumer that asks for `RVB-4471902` must get back `EID-00913338`, not a second identity.
Three things go wrong at this edge. The Match Candidate Store row can sit at `CANDIDATE_AUTO` indefinitely if the adapter's target session expires — staging looks clean, but the EMPI Identity Store never learns about `RVB-4471902`, and Riverbend's record is invisible to anyone querying by `EID-00913338`. Target-side defaults can quietly rewrite what you sent: a null middle initial materializing as an empty-string or a Riverbend gender enum the store doesn't recognize being coerced to a default value means the identity you inspect afterward doesn't match the payload you delivered. And redelivery of `RVB-2024-118773` after an interface restart can produce two candidate rows and two adapter invocations racing on the same identity, leaving the link recorded on the MRN side but not reflected on `EID-00913338`. To test, replay `RVB-2024-118773` and then read the EMPI Identity Store directly — assert `EID-00913338` lists exactly two MRNs, `CTH-220814` and `RVB-4471902`, with no third identity holding `RVB-4471902`; poll for that state with a hard timeout that fails rather than passing on the `CANDIDATE_AUTO` row existing. Then field-by-field compare the stored demographics against what was sent, specifically asserting the middle initial and gender survived unmodified rather than picking up a target default.
CP1 — Event Routing & Execution
Registration System → Registration Event Feed → Probabilistic Match Engine → Match Candidate Store → Auto-Merge Adapter
CP2 — Target Delivery
EMPI Identity Store
Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.