Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / Walmart

Retail — running since 2024

Walmart — price and stock at store level.

In grocery, the chain price is a fiction. What matters is the price and the stock in the specific store a shopper walks into, which turns the unit of work from a catalogue into a matrix — thousands of stores multiplied by a tracked basket of thousands of items, refreshed before the client's morning report.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case filewalmartIn production
SourceWalmart, full store estate
ClientGrocery price intelligence
Running since2024
DeliverySnowflake, daily
Estate4 600 stores
Basket14 000 SKUs
Reach → Read → Reconcile → DeliverRecounted daily
23B

SKU-store rows a year

64M

rows a day, every store

99.0%

measured coverage

4 600

stores covered

SourceWalmart
VerticalRetail
Running since2024 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client sells price intelligence to consumer-goods manufacturers, whose questions are always local: what does this product cost in these three hundred stores, is it on shelf, and did the promotion actually run. A chain-level average answers none of that, and the previous supplier delivered chain-level averages with a regional field bolted on.

Fixed on day one

What we committed to

Every store in the estate present every day, not a sample4 600
Shelf availability distinguished from delisting and from a failed readEvery record
Promotional state captured on the day it changes, not afterDaily
Whole estate refreshed before the client's morning runOvernight

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — The matrixThe matrix

Stores × SKUs, not SKUs

Price and stock are properties of a store, so the work is not a catalogue sweep — it is a matrix. Tens of thousands of items across thousands of stores is two orders of magnitude more observations than a naive plan assumes.

What we doStore-first traversal with per-store assortment models: each store's actual range is learned and maintained, so the matrix that gets swept is the one that exists rather than the Cartesian product that does not.

02 — GeographyGeography

Catalogues bound to a location

What is even visible depends on the store context the request carries, and getting that context wrong produces a plausible file about the wrong shop.

What we doStore context is established explicitly and verified per request against known-good markers, so an observation is either attributable to a specific store or it is discarded.

03 — PromotionsPromotions

A calendar that turns over at night

Promotional prices, multibuys and loyalty prices appear and vanish on their own schedule, and a feed that reads them a day late reports the wrong price on precisely the days the client is asked about.

What we doThe promotional calendar is swept on its own cadence, ahead of the main refresh, so what lands in the morning file is what is on the shelf that morning rather than what was there yesterday.

04 — AbsenceAbsence

Three different kinds of missing

Out of stock, not ranged in this store, and we failed to read it are three completely different facts, and collapsing them into a null is the most common defect in retail data.

What we doEach is classified separately with the evidence retained, and the classification is audited weekly against a manual check in a sample of stores.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Learn each store's real range

A store's assortment is maintained as a model and updated as it drifts, so the sweep covers what is actually ranged there. Sweeping the full catalogue against every store is the naive plan that makes this problem look impossible.

02

Attribute or discard

Every observation carries proof of which store it came from. An unattributable read is thrown away rather than delivered, because a price attached to the wrong shop is worse than no price at all.

03

Read promotions on the day they turn

The promotional calendar is swept on its own cadence, ahead of the main refresh, so the file the client gets in the morning reflects what is on the shelf that morning.

04

Classify absence explicitly

Out of stock, not ranged and read failure are three fields, not one null. The client's models depend on the difference, and the difference is audited against manual checks each week.

05

Finish before the morning run

The whole estate is scheduled to complete before the client's reporting window, with the tail of slow stores prioritised early rather than left to the end.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

SKU-store rows a day64M
Requests a day4.2M
Stores swept a day4 600
Promotional changes captured a day310k
Reads discarded as unattributable0.3%

Fourteen thousand tracked items against 4 600 stores is 64 million rows a day — 1.9 billion a month, 23 billion a year, 44 billion since 2024. They come from 4.2 million requests, 49 a second, because a store's category page carries fifteen items at once. Sweeping each store's entire assortment rather than the tracked basket would be 550 million rows a day; it is affordable, and it is the wrong thing to buy.

Headline

The four that are contractual

Rows a year
23B
Stores
4 600
Coverage
99.0 %
Attribution accuracy
99.9 %

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Granularity

Chain average with a region field

Every store, every day

Availability signal

Absent

Out of stock, delisted and read failure separated

Promotion capture

Up to a week late

Same day

Client questions answerable

Regional

Store-level, including single-store queries

The client's manufacturers had been asking store-level questions for years and getting regional answers. The product changed category the month the granularity did.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.