← All posts

Aug 30, 2026

Streaming Lakehouse Analytics Platform

How it works

The tempting places to test this pipeline are its two visible ends: the Orders API, where a request either validates or doesn't, and the BI layer, where a KPI either renders the number a stakeholder expects or doesn't. Both are cheap to assert against and both are misleading. The Orders API is a synchronous contract — one caller, one response, one schema — and the BI layer is a read surface that will happily display a confident number computed from whatever the warehouse handed it. Neither of them is where meaning changes. Meaning changes in the microservice aggregation layer, which composes an order into something wider than the order itself, and again at the event queue, which turns that composed object into an asynchronous fact with no caller waiting for it, no synchronous failure path, and no natural place for a request-scoped assertion to live. Everything downstream — the lake, the warehouse, the KPI — is a faithful transformation of whatever crossed that boundary. Faithful, not correct.

So a QA engineer's first attention belongs at the aggregation-to-queue seam and at the queue-to-lake landing, because those are the only points in the flow where a defect can be introduced silently and then be reinforced by every stage after it. An aggregation that joins the wrong context, or a queue that delivers the same event twice, or a lake that lands a partial batch, does not produce an error — it produces a plausible record. The warehouse will model that record cleanly and the BI layer will chart it without complaint, which means a KPI failure discovered at the end tells you almost nothing about where the fault entered. Testing at the endpoints gives you detection with no localization; testing at the seams gives you both. The practical consequence is that assertions belong on what the aggregation layer emits versus what the API received, and on what the lake holds versus what the queue delivered — before the warehouse gets a chance to make bad data look structured and the BI layer gets a chance to make it look authoritative.

Caveats — what breaks in practice

The failure most likely to blindside a team is small-file accumulation from high-frequency writes degrading query performance, because it is the only one on this list that arrives with no error at all. Every other mode leaves a handle to grab: duplicates on retry show up as double-counted aggregates, poison messages halt processing, TTL expiry and DLQ accumulation at least produce a queue you can inspect, schema evolution breaks a downstream consumer loudly enough that someone files a ticket. Small-file accumulation produces successful writes, successful commits, and correct data — and then queries that get slower every week. By the time anyone investigates, the cause is a write pattern that has been "working" since day one.

The assumption that makes it surprising is that a healthy system reports its own unhealthiness — that if writes succeed and rows are correct, the pipeline is fine, and any real problem will surface as a failed job, a DLQ entry, or a schema exception. That assumption holds for silent gaps only partially (a missing integration event is at least detectable by reconciliation) and holds not at all here: the degradation is emergent from write frequency, not from any single operation being wrong. For QE, this means correctness tests and error-path coverage will pass indefinitely while the defect grows. Catching it requires asserting on file counts and file sizes per partition after sustained high-frequency writes, and treating query latency against a realistically aged table as a test signal rather than a performance nice-to-have — the same discipline you'd want for detecting partial partition states from concurrent writers, where the inconsistency is also brief and also throws nothing.

How to test this end to end

Picture order `ORD-8842197` moving through CP1: the Orders API accepts a partial cancellation — two of five line items voided — and the Microservice Aggregation Layer builds an `order.updated` event carrying `order_id: "ORD-8842197"`, `event_version: "2.3"`, `status: "PARTIALLY_CANCELLED"`, `order_total_cents: 41950`, `currency: "USD"`, `updated_at: "2024-06-11T14:03:22.418Z"`, and a `line_items` array of five entries each with `line_id`, `sku`, `qty`, `unit_price_cents`, and a new `cancel_reason` field that is `null` on the three surviving lines. Because the order has five line items and the layer splits oversized payloads, the event goes out as two messages — `ORD-8842197#1` with three items and `ORD-8842197#2` with two — onto the event queue, where the lakehouse writer lands both into the `orders_events` table partitioned by `event_date=2024-06-11`. From that single landing, data science queries it ad hoc, the scheduled warehouse load picks it up for BI, and the ML pipeline reads the same rows for training features — so whatever lands here is what all three see.

The things that go wrong here compound quietly. If the aggregation layer commits the cancellation but the event emit fails, `ORD-8842197` stays `CONFIRMED` in the lakehouse forever — a silent gap with no error anywhere. If the emit times out after the broker actually accepted it, the retry lands `ORD-8842197#1` twice and the order's cancelled total double-counts. If `cancel_reason` was added when the service's own model changed, a downstream consumer expecting the old column set can break on schema drift, while the `null` values on uncancelled lines are exactly the edge-case mapping that flips a total to zero or throws. And the batch split is its own hazard: if `ORD-8842197#2` is dropped or expires on TTL during consumer downtime, the lakehouse holds three of five line items and every consumer computes a wrong total from a row that looks perfectly valid. To test it, drive a partial cancellation on `ORD-8842197` and assert the lakehouse shows exactly five distinct `line_id` values summing to `41950` — no more, no fewer; then force a transient broker failure mid-emit and re-run, asserting the same five rows and total rather than ten rows or `83900`; then kill the consumer between `ORD-8842197#1` and `#2` and confirm the missing half either arrives on recovery or surfaces as a DLQ entry that actually alerts, instead of silently leaving a three-item order as the source of truth; finally replay `ORD-8842197` with `cancel_reason` present and absent to confirm both the warehouse load and the ML feature read tolerate the added column and the nulls.

Downstream of the lakehouse landing, the scheduled warehouse load picks up the `ORD-8842197` rows from `orders_events` where `event_date=2024-06-11` and writes them into the warehouse's order-fact target. All five `line_id` values arrive, the header fields map across — `status: "PARTIALLY_CANCELLED"`, `order_total_cents: 41950`, `currency: "USD"`, `updated_at: "2024-06-11T14:03:22.418Z"` — and the two voided lines carry their `cancel_reason` while the three surviving lines carry `null`. The load reports success. But this is the one branch with a second hop: the warehouse target accepting the record is not the same event as BI showing it, so `ORD-8842197` can be delivered-and-staged in the warehouse while the dashboard an analyst is looking at still shows the pre-cancellation total, and nobody upstream of that gap sees anything wrong.

Three things bite here specifically. The staged rows for `ORD-8842197` can be accepted and then never promoted — the load job is green, the lakehouse is right, and the fact table simply never surfaces the partial cancellation, which reads as a stale order rather than a failure. Target-side defaults can quietly rewrite what landed: the `null` `cancel_reason` on the three surviving lines may be defaulted to an empty string or a placeholder code, so the warehouse now claims all five lines were cancelled and BI's cancelled-revenue number diverges from what data science and ML compute off the same lakehouse rows. And if the upstream redelivery of `ORD-8842197#1` ever produced a second copy in the landing zone, the load stages those three line items twice and `order_total_cents` inflates past `41950` in the fact table only, even though the lakehouse itself is clean. To test it, run the scheduled load for `event_date=2024-06-11` and assert not just that `ORD-8842197` was accepted but that its five rows are promoted and queryable in the fact table, then poll the BI layer until the refreshed total reads `41950` and assert on that stale window explicitly rather than assuming the load's success means the dashboard is current; separately, assert the three surviving lines still read `null` for `cancel_reason` after the write — not `""`, not `"NONE"` — and cross-check the warehouse's cancelled-line count for `ORD-8842197` against the same count taken straight from the lakehouse; finally, deliberately stage a duplicate `ORD-8842197#1` and confirm the load either dedupes on `line_id` or fails loudly, so BI never reports a total above `41950`.

At CP3 the fact-table rows for `ORD-8842197` are handed to the BI layer that builds the Insights & KPIs surface, and the shape of what crosses that boundary matters more than the row count. The five `line_id` rows carry `status: "PARTIALLY_CANCELLED"`, `order_total_cents: 41950`, `currency: "USD"`, `updated_at: "2024-06-11T14:03:22.418Z"`, and per-line `cancel_reason` — `"CUSTOMER_REQUEST"` on `ORD-8842197#4` and `#5`, `null` on `#1`, `#2`, `#3` — and the BI layer rolls these into a cancelled-revenue KPI and an order-status breakdown tile. The dashboard an analyst opens should read one order, total `41950`, two cancelled lines. The load's green status says the warehouse target accepted the rows; it says nothing about whether the BI layer read every column it needed or aggregated them exactly once.

Two things go wrong at this specific hop. If the BI layer's expected schema no longer matches what the fact table exposes — say `cancel_reason` was added or renamed after the semantic model was defined — the column is silently dropped rather than rejected, and `ORD-8842197` renders as a plain `PARTIALLY_CANCELLED` order with a cancelled-revenue contribution of zero, because the reason column that drove the KPI simply isn't in the extract. And if the BI refresh is replayed — a manual retrigger, or a scheduled refresh overlapping the one already in flight for `event_date=2024-06-11` — the aggregation can consume `ORD-8842197`'s rows twice and the cancelled-revenue KPI double-counts `#4` and `#5` while the order-count tile reports two orders where there is one, all off a fact table that is perfectly correct. To test it, assert on the columns crossing into BI, not just the rows: query the BI semantic layer for `ORD-8842197` and confirm `cancel_reason` is present and populated on `#4` and `#5`, failing the test if the field is absent rather than letting a missing column read as an empty value. Then trigger the refresh for `event_date=2024-06-11` twice — once normally, once as an immediate replay — and assert the cancelled-revenue KPI and the distinct-order count for `ORD-8842197` are identical after both runs, with the total still `41950` and exactly two cancelled lines, so a replayed refresh cannot inflate a KPI that the fact table never inflated.

CP1 — Capture & Transform

Orders API → Microservice Aggregation Layer → Event Queue → Data Lake / Lakehouse

CP2 — Target Delivery

Data Warehouse

CP3 — Sink Delivery

Insights & KPIs (BI Layer)

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.