Aug 28, 2026
CQRS + Event Sourcing Integration
How it works
The obvious endpoints of this flow are the least interesting places to spend test effort. The Command API and Command Validator sit in front of an append-only log, which means their job is narrow and their failure mode is loud: a bad command is either rejected by the validator or it becomes an event, and once it becomes an event it is a permanent fact in the Event Store. Likewise the Query API is a thin surface over Read Store A — it returns what the store holds, nothing more. Neither end owns the hard problem. The hard problem lives between the Event Store and Read Store A, in the stretch where the Event Publisher hands events to Projection Builder A and that builder decides what the read model becomes. That is the only point in the chain where an authoritative, immutable log is translated into a mutable derived shape, and translation is where meaning gets lost.
That translation is also the only asymmetric step in the flow. Everything upstream of the Event Store is append-only and therefore recoverable by definition — the log still holds the truth. Everything downstream of Projection Builder A is a consequence of how the projection interpreted that truth, and a wrong interpretation produces a Read Store A that is internally consistent, serves clean responses through the Query API, and is silently wrong. A QA engineer testing only the endpoints will see a command accepted and a query answered and conclude the system works, because both endpoints are behaving exactly as designed. The tests that actually matter are the ones that treat the Event Store as the oracle and interrogate whether Read Store A is a faithful function of it: replaying the same log through Projection Builder A and comparing outcomes, checking what happens when the Event Publisher delivers events out of order or more than once, and confirming that the Query API's answers can always be derived back to events in the log. Test the seam, not the doors.
Caveats — what breaks in practice
The failure most likely to blindside a team is the new event type that matches no subscription — black-holed. Every other item on the list eventually announces itself: validation rejects, throttling slows, timeouts time out, poison messages halt processing, consumer lag grows, retry storms spike. Even dead-lettering on delivery failure without alerting leaves a queue of evidence someone can go count. A black-holed event type leaves nothing. The publisher gets a successful write to the event store, the stream length grows, and no consumer errors because no consumer was ever asked to do anything. It sits alongside its close cousin, the subscription filter mismatch that silently discards a type, and both are dangerous for the same reason: the absence of processing is indistinguishable from the absence of work.
The assumption that makes it surprising is that a successful append to the event log implies the event will reach the read side — that publishing and consuming are one act rather than two independently configured ones. That assumption is reasonable, because in CQRS + Event Sourcing the write path genuinely is the source of truth, and it genuinely did its job. What it doesn't own is the subscription topology, which is where the event actually dies. The testing consequence is direct: asserting that an event was emitted is not a test, and neither is asserting the aggregate's state. Every new event type needs a test that names the subscriptions expected to receive it and fails when the match count is zero — the same assertion that catches a filter mismatch, and the same assertion that catches schema drift after a migration breaking a downstream reader silently. Treat "no consumer claimed this event" as a hard failure in the pipeline, not a condition to be discovered later by a stale read model that nobody can explain.
How to test this end to end
Picture a command `SubmitPolicyEndorsement` arriving at the Command API with `commandId: "cmd-9f3a-7742"`, `policyId: "POL-2201884"`, `effectiveDate: "2024-02-29"`, `endorsementType: "COVERAGE_INCREASE"`, `premiumDelta: -0.00`, and a `lineItems` array of 340 entries, one of which carries `coverageCode: null` and `limitAmount: 25000000`. The Command Validator is the only gate between this payload and the append to the event store, so everything it lets through becomes immutable history and everything it rejects never happened at all. The realistic failure shapes cluster here: the leap-day `effectiveDate` or the signed-zero `premiumDelta` trips an edge-case rejection; the null `coverageCode` maps to an unexpected enum default rather than failing loudly; a 340-line batch gets split for processing and line 171 is either dropped or written twice; the oversized batch times out mid-append so the API returns 202 for `cmd-9f3a-7742` while nothing was actually persisted; and if that null-coverage line is treated as a poison message, the whole command stream stalls behind it. The last one is the nastiest, because both read projections are derived and disposable — a projection you can rebuild, but an event that was never appended is simply gone, and no replay will bring it back.
To test it, submit `cmd-9f3a-7742` exactly as described and then assert against the event store directly, not the read side, because the projection lag is expected and will mask a missing append. Confirm 340 distinct `lineItemId` values landed with no duplicates, then resubmit the identical `commandId` and confirm a second append does not occur. Vary one field at a time off the same record — swap `effectiveDate` to `2024-02-28` and `2023-02-29`, flip `premiumDelta` between `-0.00`, `0.00`, and `null`, set `coverageCode` to null, empty string, and an unknown value like `"XZ99"` — and record for each whether the validator rejects with a specific reason or silently coerces. Push the same command at 1,000 in flight to see whether throttling returns a retryable rejection or a success response with no corresponding event, and finally inject the null-`coverageCode` line into a batch and verify that a well-formed command submitted immediately after it still gets appended rather than queuing behind the poison record.
Once the Command Validator lets `cmd-9f3a-7742` through, CP2 picks up where the append lands: 340 `PolicyEndorsementLineAppended` events plus one `PolicyEndorsementSubmitted` event committed to the append-only log for stream `POL-2201884`, then handed to the Event Publisher for delivery to the subscribers that feed the two projections. The partition key is the natural suspect — if every event for this command keys on `policyId: "POL-2201884"`, all 341 events land on one partition and that single 340-line batch hot-spots it while other policies' partitions sit idle, so consumer lag climbs on exactly the stream you care about. Subscription filters are the quieter risk: if the projection subscribers filter on `endorsementType: "COVERAGE_INCREASE"` but the `premiumDelta: -0.00` line caused the writer to emit the batch under a variant type, or if the null-`coverageCode` line was written as a distinct event type that no subscription matches, those events are appended and durable but black-holed — never delivered, so the projections look plausible and simply omit part of the endorsement. Delivery failures compound this: if line 171 dead-letters with no alert, and retention expires before anyone drains it, the event still exists in the log (rebuildable) but the operational trail that it was never delivered is gone.
Test it by appending `cmd-9f3a-7742` and then reading the log for stream `POL-2201884` to confirm all 341 events committed, and separately counting what each subscriber actually received — the gap between those two numbers is the whole point of this checkpoint. Assert the partition assignment for the 341 events and check whether they spread or collapse onto one partition, then watch consumer lag on that partition while the batch drains. Enumerate every subscription's filter against the exact event types emitted for `cmd-9f3a-7742`, including whatever the null-`coverageCode` line produced, and fail the test if any emitted type matches zero subscriptions. Then force a delivery failure on the line-171 event and confirm it reaches the dead-letter destination *and* raises an alert rather than sitting silently; finally, stop a subscriber past the retention window, restart it, and verify the endorsement events for `POL-2201884` are still consumable — and if they are not, confirm that rebuilding the projection from the log restores them, since that is the only recovery this pattern gives you.
Projection Builder A picks up the 341 events for `cmd-9f3a-7742` and folds them into the endorsement summary row for `POL-2201884` in Read Store A — one upsert of the header from `PolicyEndorsementSubmitted` (status, effective date, total premium delta) and 340 line rows keyed on `policyId` plus `lineSeq`. The mapping is per-event-subtype: the ordinary `PolicyEndorsementLineAppended` events map `coverageCode` → `read_line.coverage_code` and `premiumDelta` → `read_line.premium_delta_cents`, but line 171's null `coverageCode` and the `premiumDelta: -0.00` line are exactly the subtypes where the builder's mapping is thinnest. A null coverage code written as a variant subtype can drop into a default branch that writes the row with `coverage_code = ''`, and `-0.00` can land as `0` — the projection then looks complete (341 events consumed, 340 line rows present) while quietly misrepresenting two of them. Meanwhile the whole 340-line batch arrives as one burst against Read Store A: if the builder opens a transaction per event and holds connections from a small pool, the batch can exhaust the pool and stall writes, and if it opens one long transaction for all 340 it holds locks on the `POL-2201884` rows long enough that a concurrent redelivery of line 171 races the same row. Add replication lag and a client querying the endorsement immediately after `cmd-9f3a-7742` reads a replica that has the header but not the lines.
Test it by replaying the exact 341 events for `POL-2201884` into Projection Builder A and asserting field-by-field, not row-count-only: `lineSeq 171` must carry the null coverage code as an explicit null (or whatever the contract says) rather than an empty string, and the `-0.00` line must land with the sign and scale the source event carried. Enumerate every event subtype emitted for `cmd-9f3a-7742` against the builder's mapping table and fail if any subtype falls to a default branch. Drive the batch through with the connection pool set to its production size and assert no timeouts leak into other streams, then deliberately redeliver line 171 concurrently with the original and assert the resulting row is identical either way. Expire the target credential mid-batch and confirm the builder retries with backoff and a cap rather than storming Read Store A. Apply a schema migration to `read_line` and re-run the same replay to prove the builder fails loudly instead of silently dropping a column. Finally, write `cmd-9f3a-7742` and immediately query a replica: assert the client sees a defined stale-or-absent state rather than a half-built endorsement, and prove the recovery path by truncating Read Store A's rows for `POL-2201884` and rebuilding from the log until the projection matches the assertions above exactly.
The delivery leg here is the Query API serving `GET /endorsements/POL-2201884` off Read Store A moments after `cmd-9f3a-7742` commits. The response is assembled from exactly what the builder wrote: a header block carrying `status: "SUBMITTED"`, `effectiveDate: "2024-11-01"`, `totalPremiumDeltaCents: -418200`, and a `lines` array of 340 entries each shaped `{ lineSeq, coverageCode, premiumDeltaCents }`. Line 171 comes back as whatever the projection actually stored — so if the default branch wrote `coverage_code = ''`, the API faithfully serves `"coverageCode": ""`, and the `-0.00` line serves `"premiumDeltaCents": 0`. Neither is an error anywhere in the chain; the caller just gets a clean 200 with two silently wrong lines, and a `count: 340` that agrees with itself.
Three things bite at this boundary. The API's own response validation or serializer can reject or mangle the edge-case values — a schema that declares `coverageCode` as a non-empty string will 500 or strip line 171 rather than surface the null, and a scale-normalizing serializer erases the sign on the `-0.00` line before it ever reaches the client. Under the 340-line burst the read path throttles: if the endpoint pages or fans out per line, concurrent callers hitting `POL-2201884` while the builder is still writing can trip rate limits and get 429s that the client retries into more load. And because reads hit a replica, the API can return 200 with the header present and `lines: []` — a success response for a record that isn't fully persisted where the caller is reading. To test, replay the 341 events, then query `POL-2201884` through the Query API and assert the response body field-by-field against the source events: `lineSeq 171` must serialize the null coverage code exactly as the contract specifies rather than `""`, and the `-0.00` line must keep its sign and scale after serialization. Hit the endpoint concurrently at production rate limits during the 340-line write burst and assert every response is either a complete endorsement or an explicit stale/not-ready signal, never a 200 with a partial `lines` array; assert 429s carry retry guidance and don't cascade. Then truncate Read Store A's rows for `POL-2201884`, rebuild from the log, and re-run the same field-by-field assertion to prove the API returns the identical body after a rebuild as before it.
CP1 — Command Capture & Validation
Command API → Command Validator
CP2 — Event/Message Persistence
Event Store (append-only log) → Event Publisher
CP3 — Event Routing & Execution
Projection Builder A (read model 1) → Read Store A
CP4 — Target Delivery
Query API
Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.