Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / Amazon

Marketplaces — running since 2023

Amazon — 180M SKUs, a full pass every week.

A repricing engine that runs hourly against week-old data is an expensive way to lose money slowly. The client needed the whole catalogue they compete in — not a watchlist — with offers and buy-box state, at a cost per million rows that left room for a business. New competitors and new listings enter the catalogue automatically, the day they appear.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case fileamazonIn production
SourceAmazon, regional storefronts
ClientMarketplace repricing platform
Running since2023
DeliveryS3 Parquet + Postgres
Catalogue180M SKUs
Full pass7 days
Reach → Read → Reconcile → DeliverRecounted daily
40B

SKU records since 2023

16B

a year

98.4%

measured coverage

180M

SKUs in the catalogue

SourceAmazon
VerticalMarketplaces
Running since2023 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client reprices inventory for thousands of sellers, which means their competitive picture has to include the whole catalogue those sellers compete in — not the few thousand ASINs somebody thought to watch. Their previous system tracked a watchlist, so every new competitor arrived invisible and stayed invisible until a human noticed the margin.

Fixed on day one

What we committed to

Full refresh of the tracked catalogue, end to endUnder 24 h
Offer and buy-box state captured alongside the price, not inferredEvery record
Cost per million rows, fixed before the buildCapped
New competitors discovered automatically, not added by handContinuous

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — ScaleScale

A catalogue that does not fit in a plan

One hundred and eighty million tracked SKUs is 16 billion records a year, and the interesting fields are the ones that require more than a listing page. Touching every SKU every day would be seven times the traffic for no extra signal.

What we doRefresh is incremental and priority-driven: volatility, competitive relevance and the client's own exposure decide what gets touched this cycle. A uniform sweep of this catalogue is neither affordable nor useful.

02 — VariationsVariations

Half the catalogue is hidden

Items live in variation trees — size, colour, pack — where the parent shows one price and the children have their own. A collector that reads parents reports a catalogue that looks complete and is missing most of the actual SKUs.

What we doThe catalogue is modelled as a graph rather than a list. Variation children are enumerated explicitly and reconciled against the parent, and the count is checked against an independent sample.

03 — Buy boxBuy box

A price that belongs to a moment

The winning offer rotates between sellers through the day, so a single daily observation of a price is a sample of a distribution presented as a fact.

What we doBuy-box state is sampled through the day for the segment that matters to the client and stored as observations with timestamps, so their engine can reason about the distribution instead of a single point.

04 — PersonalisationPersonalisation

What you are shown is about you

Storefront, session history and locale all change what appears and at what price, which makes uncontrolled observations mutually incomparable.

What we doObserver identity is fixed per storefront and segment, so a movement in the series is a movement in the market and not a movement in who was looking.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Priority beats uniformity

Every cycle the pipeline scores what to refresh: how volatile the item has been, how much of the client's revenue touches it, how long since it was last seen. Uniform refresh spends most of its budget on items that have not changed in a month.

02

Model the catalogue as a graph

Parents, variation children and offers are separate nodes with explicit edges, so completeness can actually be checked. A flat list of items cannot tell you what it is missing.

03

Sample the buy box, do not snapshot it

For the competitive segment the client cares about, the winning offer is observed repeatedly through the day and delivered as a series. A single daily price is the most common way this data is quietly wrong.

04

Discover, do not curate

New competitors and new SKUs enter the picture through category and search traversal rather than through somebody adding them to a list. The watchlist model fails at exactly the moment a new competitor matters.

05

Hold the cost line

Cost per million rows was agreed before anything was built and is reported weekly. At this volume it is the constraint that shapes every other decision, from the request path to how much of the catalogue is refreshed each cycle.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

SKU records a day45M
SKUs touched a day26M
Requests a day8.6M
Offers captured a day9.4M
Requests refused by the source1.1%

One hundred and eighty million SKUs cannot be touched every day and do not need to be. Twenty-six million a day is a complete pass every seven days, with volatility and the client's own exposure deciding what jumps the queue — a contested SKU is read several times a day, a dormant one every few weeks. A listing page returns twenty items at once and a detail page one, and the mix works out at about three SKUs per request, so 26 million SKUs cost 8.6 million requests, 100 a second. Each read writes buy-box, offer and stock rows: 45 million records a day, 1.35 billion a month, 16 billion a year, 40 billion since 2023.

Headline

The four that are contractual

Records since 2023
40B
Coverage
98.4 %
SKUs tracked
180M
Cost / 1M rows
−54 %

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Catalogue in view

A curated watchlist of 40k

The full 180M competitive catalogue

New competitors

Noticed by a human, eventually

Discovered automatically

Price observations per day

One, per item

Buy-box series through the day

Repricing decisions on stale data

Routine

Under 3%

The change the client noticed first was not the volume. It was that competitors stopped appearing out of nowhere three weeks after they had started eating the margin.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.