Retail & consumer
Cutting stockouts across 240 stores with a forecast the buyers trust
A hierarchical demand forecast, a promo-aware feature store and a replenishment workflow the category team can override — because a model nobody overrides is a model nobody uses.
- Stockouts on A-class SKUs
- cut by roughly a third
- Forecast accuracy (WAPE)
- mid-thirties → high teens
- Buyer overrides accepted
- zero → most plans
- Sector
- Multi-category retail, India
- Estate
- 240 stores, 3 DCs, ~28,000 SKUs
- Engagement
- Discovery sprint, then delivery pod
Stack
- LightGBM
- Feast
- dbt
- Snowflake
Practices involved
Discuss a similar problemThe situation
The chain ran replenishment from a weekly spreadsheet cycle. Category managers pulled sales from the ERP, adjusted by memory for festivals and promotions, and emailed order quantities to three distribution centres. The process worked when the business had 60 stores. At 240, the same three people were making roughly 28,000 decisions a week, so most SKUs were simply reordered at last week's number.
The symptoms were the ones you would expect: fast movers went out of stock in the last four days of the month, slow movers accumulated in the DC, and the two problems were invisible to each other because nobody reported them together.
The constraint
An earlier forecasting pilot had been rejected. It was more accurate than the spreadsheet on paper, but it produced numbers the category team could not explain to their own management, and it had no way to absorb the thing they knew and it did not — a competitor opening opposite Store 118, a distributor strike, a school holiday shifted by a week.
So the requirement was not really accuracy. It was accuracy the buyers would accept ownership of.
What we built
A feature store that knows what a promotion is
Promotions lived in three systems and a shared drive. We modelled them once — mechanic, depth, participating SKUs, store list, date range — and made that the single input to both training and serving. Festival calendars, local holidays, weather and competitor-opening flags landed in the same place. Training-serving skew disappears when both paths read the same table.
A hierarchical forecast, reconciled
Separate models at SKU-store, SKU-cluster and category-region level, then reconciliation so that the sum of the store forecasts equals the number the category head is held to. Gradient-boosted models on intermittent-demand SKUs, a seasonal statistical model on the long-tail slow movers where there was not enough signal to justify anything heavier.
An override that teaches
The replenishment screen shows the recommendation, the three factors that moved it most, and a field for the buyer to change it with a reason code. Every override is stored. Two things happen with that data: recurring reason codes become candidate features, and the acceptance rate per buyer became the metric we managed the rollout by.
A rollout designed to be reversed
Six weeks of shadow mode with no orders placed, then one region, then three, then all. At every stage the previous process stayed live and switchable within a day.
What changed
The number the client cares about is not the WAPE improvement, it is that most replenishment plans now go through with a buyer adjustment rather than a wholesale rejection. The forecast became a starting position instead of an opinion competing with theirs.
What we would do differently
We spent four weeks building a clean master-data mapping between ERP item codes and the point-of-sale catalogue before we understood how often the mapping changed. It changed weekly. We should have built the reconciliation job first and the model second; instead we rebuilt the mapping twice.
Outcomes
- Stockouts on A-class SKUs
- cut by roughly a third
- Forecast accuracy (WAPE)
- mid-thirties → high teens
- Buyer overrides accepted
- zero → most plans
Client identity withheld under a mutual NDA. Figures are illustrative — rounded and directional, meant to show the shape of the change rather than an audited result. We will walk through the real numbers, how they were measured, and put you in touch with a reference, under NDA on a call.
More work
Other engagements.
Case studiesEnergy analytics across 14 buildings and four BMS vendors
One data plane over BACnet, Modbus and a proprietary head-end, with fault detection that tells a facilities manager which asset to look at before the complaint arrives.
Read the case studyTurning 40,000 support calls a month into a product backlog
Diarised transcription, an LLM extraction pipeline with a graded evaluation set, and a review console where quality leads correct the model in place.
Read the case studyTeaching a voice agent to negotiate a truck booking, not just take one
An outbound agent that calls transporters directly, holds a rate band instead of reading a script, and hands off to a human the moment a call needs one — built for a market that runs on phone calls, not portals.
Read the case studyNext step
Tell us what you're trying to ship.
Send the brief, the RFP, or three messy sentences about the problem. You get a written point of view from an architect within two working days — not a sales deck.