Aug 20, 2026
Change Data Capture (CDC) Replication
How it works
The tempting places to look are the two ends: does the source row change, and does the warehouse row match. Both are cheap to assert and both will pass long after the pipeline has quietly gone wrong, because the endpoints only tell you about state, and CDC is a flow about *events*. The real complexity sits in the middle band — the Schema Registry / Transform stage and the Fan-out Router — because that is the only place in this chain where a single input becomes something other than a single faithful output. The CDC Connector's job is comparatively mechanical: observe the source and emit change events to the Change Event Topic. The Sink Connector's job is likewise mechanical: take what arrives and land it in the Analytics Warehouse. Neither of them decides anything. The Transform stage decides what a change event *means* under a given schema version, and the Fan-out Router decides how many places that meaning travels to. Decisions are where defects live.
For a QA engineer this reorders the test plan. Endpoint reconciliation between Source Database and Analytics Warehouse is a lagging indicator — it tells you something broke, not which decision broke it, and it cannot distinguish a transform that dropped a field from a router that never delivered the event at all. Test the Change Event Topic as a first-class artifact: what the connector actually emitted, in what shape, before anything interprets it. Then test the Transform against schema evolution deliberately, because the Schema Registry exists precisely because schemas change, and a stage that exists to absorb change is a stage whose behavior under change is the thing worth exercising. Then test fan-out as a multiplicity problem — one event in, correct set of destinations out — rather than as a delivery problem. Do that, and the warehouse comparison becomes confirmation rather than investigation.
Caveats — what breaks in practice
The failure most likely to blindside a CDC team is schema drift after a migration silently breaking a downstream reader — specifically its quieter cousin, the silent schema mismatch that drops columns. Pick it over the loud candidates: connection pool exhaustion announces itself as timeouts elsewhere in the pipeline, oversized batches announce themselves as timeouts, poison messages announce themselves by halting processing outright. Even replication lag has a familiar shape — someone reads a replica, gets a stale result, and the staleness is at least the kind of wrongness people know to look for. Schema drift produces no timeout, no halt, no dead letter. The pipeline keeps running at full throughput, the consumer keeps committing, and the only evidence is a column that stopped arriving.
The assumption that makes it surprising is that a migration which succeeded upstream is a migration that arrived downstream — that because CDC replicates changes faithfully, structural changes travel with the data changes and any incompatibility would break something visibly. It's a reasonable belief: the entire premise of CDC is fidelity to the source. But the reader has its own picture of the schema, and when those pictures diverge, the mismatch resolves by dropping what doesn't fit rather than by raising an error. For testing, this means migration scenarios can't be validated by confirming the pipeline still flows — flow is exactly what's preserved. You need assertions on field-level completeness across a schema change, including edge-case values like nulls and unexpected enums where mapping errors hide, and you need a test that a newly added or renamed column actually lands rather than merely that the consumer stayed up.
How to test this end to end
Testing a CDC pipeline checkpoint-by-checkpoint in sequence feels orderly, but it spends your earliest and best-funded test cycles on the part of the system least likely to betray you. The Analytics Warehouse Sync path carries a risk score of 19, and that risk is not distributed evenly across the three checkpoints — it concentrates in two components, the CDC Connector and the Schema Registry / Transform. Those two failure modes deserve to be attacked first, before the sink is exercised at all, because they are the failures that produce no error and no alarm. A silently lost event and a silently mismapped field both leave a pipeline that reports success end to end. Order your testing by where the silence lives, not by where the data flows.
Start at CP1, Capture and Buffer, where the source database hands changes to the CDC Connector and the connector writes to the change event topic. Correct here means writes commit transactionally and are immediately visible to any reader — there is no deferred posting job, no eventual-consistency window to blame a missing row on. That property is exactly what makes this checkpoint the right place to begin, because it removes ambiguity: if a record is committed at the source and absent downstream, the loss is real and it happened after the commit. Validate it by writing a known record and immediately reading it back on a separate connection, which proves visibility is not an artifact of your own session, and by running a representative query load while watching connection-pool saturation, which tells you whether the connector's fetch behavior degrades under realistic pressure. The specific risk to hunt is the connector's fetch failing after the notification has already been acknowledged, which loses the event silently. That is why the separate-connection read-back matters more than it looks: it establishes ground truth at the source that you can later reconcile against, and it is the only way to distinguish "never captured" from "captured and dropped." Test this first because every downstream assertion you write is worthless if you cannot trust that the event existed and was acknowledged.
CP2, Transform and Publish, is the second priority and the second concentration of risk. Correct means every input record maps to a canonical output with all required fields, and no records are dropped or duplicated when a batch is split for fan-out. Validate it by feeding a crafted payload with known field values and diffing the canonical output field by field — not spot-checking, field by field, because the named risk is mapping errors on edge-case field values such as nulls and unexpected enums, and those are precisely the fields a summary comparison skips over. A null that maps to a default, or an unrecognized enum that falls through to a catch-all, produces a well-formed record that passes every structural check and is wrong. The second validation, sending an oversized batch at ten times normal volume and watching duration against the timeout limit, addresses the other half of this checkpoint: the batch split. If the split times out, records are dropped or duplicated, and the count discrepancy is small enough to be mistaken for timing. Both of these failures reach the warehouse looking like data. Neither raises an error. That is the argument for putting CP2 ahead of sink delivery.
CP3, Sink Delivery into the Analytics Warehouse, is genuinely important and genuinely last. Correct means dashboards and reports reflect ingested events within the stated latency window. Validate it by injecting a marked test event and finding it in the analytics layer, and by replaying a batch and verifying counts do not double. Both of these are strong tests, and both are strong tests of the wrong thing if run first. A marked test event is, by construction, a clean payload with no nulls and no unexpected enums, so it will sail through a transform that is quietly mangling edge cases in production traffic. And a replay-and-count check confirms idempotency at the sink, which tells you nothing about events the connector never delivered because it acknowledged the notification and then failed to fetch. Run CP3 after CP1 and CP2 and its results become diagnostic: a marked event that arrives proves the whole path is wired, and counts that hold under replay prove the sink is safe to retry against. Run it before, and a green result is just an untested pipeline that happened to be handed easy input.
The practical consequence for a QE team is that the test plan should not mirror the architecture diagram. Build the reconciliation harness at CP1 first so you have an authoritative record of what was committed and acknowledged. Then build the field-by-field canonical diff at CP2, weighted toward nulls and unexpected enums, and the oversized-batch timing test against the split. Only then inject the marked event and run the replay against the warehouse. Ordered this way, a CP3 failure is informative rather than confusing, because you have already eliminated the two silent failure modes that carry the risk on this path.
CP1 — Capture & Buffer
Source Database → CDC Connector → Change Event Topic
CP2 — Transform & Publish
Schema Registry / Transform → Fan-out Router
CP3 — Sink Delivery
Sink Connector — Analytics Warehouse
Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.