Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / eBay

Marketplaces — running since 2021

eBay — sold prices, not asking prices.

For anything second-hand, the asking price is an opinion and the sold price is a fact. The client needed the facts — at category scale, across five markets, with enough seller and condition detail to explain why two identical items sold eleven days and forty per cent apart. A sale can only be caught if the listing was already being watched, so every new listing enters the scan the day it is posted.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case fileebayIn production
SourceeBay, 5 markets
ClientResale valuation platform
Running since2021
DeliverySnowflake, daily
Listings watched340M
Markets5
Reach → Read → Reconcile → DeliverRecounted daily
2.6B

sold records to date

620M

sold records a year

98.0%

measured coverage

64M

listings scanned a day

SourceeBay
VerticalMarketplaces
Running since2021 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client values second-hand goods for insurers and resale platforms, which means their entire model rests on realised prices rather than listings. Sold data is thinner, shorter-lived and harder to collect than active listings, which is exactly why most competitors quietly model from asking prices and call it a valuation.

Fixed on day one

What we committed to

Sold observations captured before they age out of visibilityDaily
Condition, seller rating and shipping retained per observationEvery record
Category taxonomy mapped to the client's own modelMapped
Five markets in one schema with local currency retained5 markets

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — WindowWindow

Sold data expires

Realised prices are visible for a limited window and then gone. Miss the window and that transaction never existed, which biases a valuation model towards whatever remains visible.

What we doCollection is scheduled against the visibility window per category, with the busiest categories swept daily and a completeness check that specifically looks for observations lost to expiry.

02 — ConditionCondition

Identical is not identical

Two of the same item at different conditions, with different shipping and different seller ratings, are different products economically. Flattening them produces a valuation with enormous unexplained variance.

What we doCondition, shipping treatment and seller quality are captured per observation and carried into the model, so the variance is explained rather than averaged.

03 — TaxonomyTaxonomy

Categories drift

Marketplace taxonomies are reorganised regularly, and a model keyed to last year's category tree quietly stops matching.

What we doThe taxonomy is versioned with explicit mappings into the client's own model, and structural changes raise an alert rather than silently redistributing volume.

04 — SellersSellers

Volume hides behind a few accounts

A category dominated by three professional sellers behaves nothing like one with a long tail, and a price distribution alone does not show the difference.

What we doSeller identity and type are retained per observation, so the client can distinguish a market with real depth from one with three sellers and an illusion of liquidity.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Schedule against the visibility window

Categories are swept at a cadence derived from how long sold data stays visible in each, with a completeness check that measures what was lost to expiry rather than assuming nothing was.

02

Carry condition and seller quality

Every observation keeps the attributes that explain price variance. A sold price without condition is a number with no error bars.

03

Version the taxonomy

Category structure is versioned and mapped explicitly, and reorganisations raise alerts. A model silently rekeyed by somebody else's reorganisation is the quietest failure in this business.

04

Keep the seller dimension

Seller identity and type travel with the observation, so depth and concentration are visible rather than inferred from a price histogram.

05

Keep local currency

Prices stay in the market currency with the observed rate recorded. The normalised figure is derived and can be recomputed when the client changes their method.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

Listings scanned a day64M
Requests a day6.4M
Listings under watch340M
Sold records captured a day1.7M
Sales lost to expiry1.7%

Three hundred and forty million watched listings turn over on a five-day cycle — 64 million scanned a day from 6.4 million requests, 74 a second, because a results page carries ten items. That yields 1.7 million confirmed sold records a day, 51 million a month, 620 million a year, 2.6 billion since 2021. One point seven per cent of sales are lost because the visibility window closes before we reach them; publishing that figure is what lets the client's model correct for the bias instead of absorbing it.

Headline

The four that are contractual

Sold records to date
2.6B
Coverage
98.0 %
Listings watched
340M
Taxonomy versions kept
All

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Valuation basis

Asking prices, adjusted by a factor

Realised sold prices

Variance explained

Largely unexplained

Condition, shipping and seller quality

Category drift

Silent model degradation

Alerted and remapped

Market depth

Invisible

Seller concentration per category

The insurer using the client's valuations asked how sold data was obtained, how complete it was, and what happened to what was missed. All three had answers, which is not usually the case in this category.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.