← All posts

Sep 8, 2026

Demand Forecasting Pipeline

How it works

The tempting places to test this pipeline are its two endpoints: the sales history warehouse, where you can count rows and check totals, and the ERP purchase orders, where a wrong number becomes money spent on stock nobody asked for. Both are easy to assert against, and that is exactly why they are poor places to spend your first effort. The warehouse is a record of what already happened, and the purchase orders are the last thing in the chain — by the time a bad PO exists, every transformation that produced it has already run and committed. The real complexity sits in the feature pipeline, because promo, price, and seasonality are not facts read out of sales history; they are constructed interpretations of it. A promotion has to be attributed to the right SKU over the right window, a price has to be the price that was actually in effect rather than the one currently on file, and seasonality is an inference imposed on the series rather than a column pulled from it. Each of those is a judgment encoded in code, and a wrong judgment produces a feature that is perfectly well-formed, passes any schema check, and is quietly false.

The second concentration of complexity is the replenishment planning job, for a different reason: it is where a forecast stops being an estimate and becomes an instruction. The demand forecasting model consumes features and emits a prediction, and a prediction is inherently a range of plausible values — you cannot assert that one number is the correct one. But the replenishment job takes that soft output and converts it into a hard, discrete order quantity that flows into ERP as a purchase order. That conversion is deterministic logic, which means it is testable in a way the model is not, and it is also the point where a modest forecast error gets amplified or damped depending on how the planning rules treat it. So a QA engineer's attention belongs at the two deterministic transformations that bracket the model — the feature construction feeding it and the planning logic consuming it — rather than at the warehouse that merely stores history or the purchase orders that merely record the outcome. Test that features reflect the promo, price, and seasonality conditions that actually held, and test that a given forecast value produces the order quantity the planning rules say it should. Those are the places where a defect is both possible and provable.

Caveats — what breaks in practice

The failure most likely to blindside a team is the stale model that keeps serving after a retraining job fails with no alert. Compare it to its neighbors on this list: a poison message halts processing, a timeout on oversized batches throws, a job failure leaves records permanently unposted — all of these announce themselves, either as a stopped queue or a visible gap in posted records. The stale model does the opposite. The serving path stays up, latency is normal, every request returns a prediction, and the only thing that broke was an upstream job whose output nobody is checking for freshness. Training data drift and training/serving skew degrade silently too, but they degrade gradually; a stale model is a clean break in provenance that still looks perfectly healthy from every endpoint a monitor is pointed at.

The assumption that makes it surprising is that a successful response from the serving layer implies a successful pipeline behind it — that if predictions are coming back, the thing producing them is current. That is a reasonable expectation, and it holds for most of the other failures here: a batch split that loses line items or a partial posting eventually shows up as a reconciliation discrepancy, because the artifact under test is the record itself. With a model, the artifact under test is the *version*, and nothing in the request/response contract exposes it. For QA this means availability checks and output-shape assertions on the prediction endpoint are not coverage; you need an assertion on model freshness — the age or version of the artifact currently loaded — treated as a first-class test, plus an explicit failure signal on the retraining job rather than inferring its health from the fact that serving still works.

How to test this end to end

Consider the nightly feature build for week `2024-W07`, where the pipeline pulls rows from `sales_history.daily_sku_location` and joins them against the promo calendar and price master to publish a feature row per SKU-location-week. A single record looks like this: `sku_id = "SKU-4471-BLK-M"`, `location_id = "STR-0192"`, `week_start = "2024-02-12"`, `units_sold = 318`, `net_revenue = 4770.00`, `avg_unit_price = 15.00`, `promo_flag = true`, `promo_type = "BOGO"`, `promo_depth_pct = 50`, `week_of_year = 7`, `holiday_flag = false`. That row is assembled from seven daily sales rows aggregated up to a week, enriched with promo and price attributes, and published as the training/scoring feature set the ML-pipeline node consumes downstream to forecast demand for that SKU-location-week.

The dangerous thing is that every one of the known failure modes here produces a row that still looks like a row. If the price master introduces a new enum and `promo_type` arrives as `"BOGO_50_TIERED"` instead of `"BOGO"`, a mapping that only recognizes the known set may silently coerce it to null or to `promo_flag = false`, so SKU-4471-BLK-M's 318 units get presented to the model as clean baseline demand at full price — inflating the learned baseline and, through the planning layer, generating a purchase order for stock the store only ever moved under a half-off promotion. Equally, if the nightly job is replayed after a partial failure and the daily rows for 2024-02-12 through 2024-02-14 are aggregated twice, `units_sold` becomes 500-odd with no error anywhere; a schema change that drops `promo_depth_pct` entirely just yields a column of nulls the model treats as "no discount." Test this by asserting on the published feature row itself, not on job status: reconcile `units_sold` and `net_revenue` for SKU-4471-BLK-M / STR-0192 / 2024-02-12 back to the sum of its seven source daily rows and fail on any delta, then rerun the same batch and assert the value is still 318 (idempotency on replay). Add a contract check that every published column in the expected schema is present and that `promo_type` values fall in the allowed enum set, routing an unmapped value like `"BOGO_50_TIERED"` to a quarantine table with an alert rather than defaulting it; seed a deliberately malformed row into the batch to confirm it quarantines instead of halting the run, and load-test a batch sized well past a normal night's volume to confirm the split neither drops nor duplicates SKU-4471-BLK-M's line items.

Carry SKU-4471-BLK-M / STR-0192 / week_start 2024-02-12 forward past the feature set: the model consumes that row and emits a forecast for the following week — say `forecast_units = 340`, `forecast_horizon_week = "2024-02-19"`, `model_version = "demand_fc_v3.2"` — and the replenishment planning job converts it into an ERP purchase order line, `po_line_id = "PO-88213-004"`, `order_qty = 340`, against vendor lead time and on-hand. The interesting things happen in that conversion and posting. The planning job rounds to a case pack of 24, so 340 becomes 336 or 360 depending on the rounding rule, and nothing in the forecast record records which one happened; the ERP itself may apply a target-side default — a minimum order quantity of 500 on that vendor — and quietly overwrite `order_qty` so the posted PO bears no resemblance to the 340 the model asked for. Worse, if the nightly retraining job failed, `model_version` is still `demand_fc_v3.2` from three weeks ago and the forecast for SKU-4471-BLK-M is generated from a model that never saw the February BOGO pattern at all, yet the run is green end to end. And because the feature row is the model's only view of reality, a serving-time computation of `avg_unit_price` that differs from how training computed it — full price versus promo-adjusted — skews the forecast before a single PO line is ever written.

Test the posted artifact, not the job exit code. Assert that `PO-88213-004` carries a traceable link back to the forecast record for SKU-4471-BLK-M / STR-0192 / 2024-02-19 and that `order_qty` equals `forecast_units` after exactly one declared transformation — record the rounding rule and the pre- and post-rounding values on the PO line, then fail the check if the ERP-returned `order_qty` differs from what was submitted, which is what catches the vendor minimum-order-quantity override. Assert `model_version` on every forecast row is within the freshness window and alert when a retrain fails rather than letting `demand_fc_v3.2` serve on; recompute `avg_unit_price` for SKU-4471-BLK-M using the training-time code path and fail on any mismatch with the serving value, which is the concrete training/serving skew probe. For posting mechanics, replay the same planning batch and assert only one PO line exists for SKU-4471-BLK-M / STR-0192 / 2024-02-19 rather than a duplicate at 340 more units, submit a batch large enough to risk a mid-batch timeout and confirm every line either posts or rolls back with none left staged, and add a staged-state age check that alerts if `PO-88213-004` sits unposted past the SLA instead of waiting for someone to notice the store never got stock.

CP1 — Transform & Publish

Sales History Warehouse → Feature Pipeline (promo/price/seasonality)

CP2 — Target Posting & Application

Demand Forecasting Model → Replenishment Planning Job → ERP Purchase Orders

Want this level of breakdown for your own system? Match your architecture in a few questions — no confidential upload required.