Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / StubHub

Ticketing — running since 2020

StubHub — the resale market, priced against itself.

Resale prices only mean something next to the other resale prices for the same seats. The client needed the whole secondary book — every listing, its exact seats, its fees and its age — refreshed fast enough that a broker's repricing decision is made against the market as it is, not as it was this morning. Every new event and every new listing enters the book the day it appears.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case filestubhubIn production
SourceStubHub
ClientTicketing analytics platform
Running since2020
DeliveryKafka + ClickHouse
Freshness, active events6 min
Live listings9.4M
Reach → Read → Reconcile → DeliverRecounted daily
230B

listing records since 2020

59B

a year

98.8%

measured coverage

9.4M

listings tracked live

SourceStubHub
VerticalTicketing
Running since2020 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client sells pricing analytics to ticket brokers, who reprice several times a day. Their previous data came from a partner feed that arrived every four hours and omitted fees entirely, so the price a broker saw in the tool was never the price a buyer would pay — which made every comparison in the product slightly wrong in a direction nobody could predict.

Fixed on day one

What we committed to

Freshness on active events, measured at p95Under 5 min
Fee structure captured per listing, not modelled from a ruleEvery listing
Seat-level detail retained, not collapsed to a section averageRequired
Listing age and delisting events preserved for velocity analysisFull history

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — VolumeVolume

A book that turns over hourly

Listings appear and vanish continuously, so a snapshot taken once is a snapshot of something that no longer exists. Velocity — how fast inventory moves at a price — is the signal brokers actually pay for, and it only exists if you never miss a turn.

What we doEvent-level change detection drives the refresh: events whose book is moving get swept in minutes, quiet events fall back to hours. Listing appearance and disappearance are recorded as events rather than inferred from diffs.

02 — FeesFees

The price is not the price

Service fees, delivery fees and taxes are computed late in the flow and vary by listing, quantity and market. A feed that reports the headline price is reporting a number no buyer ever pays.

What we doThe fee stack is captured per listing at the quantity the client cares about and stored as separate components, so the product can show both the headline and the landed price.

03 — SeatsSeats

Section averages hide the money

Two listings in the same section can differ by a factor of three depending on row and sightline. Collapsing to a section average destroys exactly the variance a broker is trying to exploit.

What we doSeat identifiers are parsed and normalised against a venue model, so listings are comparable at row level across events and across time.

04 — IdentityIdentity

Sessions that go stale

Long-lived sessions get scored down over hours, and the failure is gradual: coverage degrades before anything is refused outright, which is the hardest kind of failure to notice.

What we doSession health is monitored as a metric in its own right, with rotation driven by observed degradation rather than by a fixed timer.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Track the book, not the page

The unit of work is an event's order book, and change detection runs at that level. Sweeping pages uniformly is how a resale feed ends up four hours stale on the events that matter.

02

Capture fees where they are computed

Fee components are read at the point in the flow where they are actually calculated, at the quantity the client models. Reconstructing them from a published rule is how most competitors get landed prices wrong.

03

Normalise seats against a venue model

Rows, sections and sightlines are mapped to a maintained venue model so a listing is comparable across events at the same venue and across venues of the same class.

04

Record appearance and disappearance

A listing that vanished is a sale or a withdrawal, and the difference is worth money. Both are emitted as events with timestamps rather than showing up as a silent absence in the next snapshot.

05

Watch session health as a metric

Degradation is measured continuously, and rotation happens when the numbers say so. A fixed timer either rotates too early, which is expensive, or too late, which quietly costs coverage.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

Listing records a day162M
Order-book fetches a day2.7M
Live listings tracked9.4M
Listing events emitted a day4.1M
Requests refused by the source0.7%

One request returns an event's entire order book, not a single listing — about sixty listings — which is the only reason this is affordable. 9 000 active events every six minutes over eighteen hours is 1.6 million fetches; the remaining 131 000 events, swept eight times a day, add 1.1 million. That is 2.7 million requests a day, 31 a second, and 162 million listing records: 4.9 billion a month, 59 billion a year, roughly 230 billion since 2020. Collecting listing by listing at this cadence would need sixty times the traffic and the source would not tolerate it.

Headline

The four that are contractual

Collected since 2020
230B records
Freshness, active events
6 min
Coverage
98.8 %
Live listings
9.4M

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Data age at repricing

Up to 4 hours

Six minutes on active events

Price shown to brokers

Headline only

Headline and landed, separately

Granularity

Section average

Row level, normalised

Sale velocity signal

Not available

Derived from listing events

Velocity turned out to be the feature brokers actually renewed for. It was not in the brief — it fell out of recording disappearances properly.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.