Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / Ticketmaster

Ticketing — running since 2019

Ticketmaster — 57 billion price records a year.

A secondary marketplace cannot price anything without knowing what the primary market is doing right now — but it only needs that for the events that are actually moving. We built the pipeline that decides which ones those are, every few minutes, across nine countries. Running continuously since 2019, and every new event that goes on sale is picked up the same day.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case fileticketmasterIn production
SourceTicketmaster
ClientSecondary ticketing marketplace
Running since2019
DeliveryKafka stream + Postgres
Freshness, hot tier3 min
Coverage SLA99% floor
Reach → Read → Reconcile → DeliverRecounted daily
210B

price records since 2019

57B

a year

99.2%

measured coverage

155M

records a day

SourceTicketmaster
VerticalTicketing
Running since2019 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client prices secondary inventory against the primary market. Before us they bought a nightly feed, which meant every pricing decision on a hot event was made against data up to fourteen hours old. What they needed was not more rows — it was the same rows, minutes old, for the events that actually move, without paying minute-level prices for the long tail that does not.

Fixed on day one

What we committed to

Freshness for volatile events, measured at p95 rather than on averageUnder 2 min
Full catalogue coverage including the long tail nobody else bothers with99% floor
Cost per thousand refreshes, fixed before the build startedCapped
Load shaped so the source is never degraded, at any hourHard rule

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — QueuesQueues

Waiting rooms

Queue systems put every client into a lobby before anything is served. A worker that simply retries burns its entire budget waiting, and the events worth watching are exactly the ones with the longest queues.

What we doA queue-aware scheduler with a budget per event: the pipeline decides which events deserve a slot in the next few minutes and which can wait for the slow sweep, instead of spending capacity uniformly.

02 — ScoringScoring

Behavioural session scoring

The source does not score what you send, it scores how you move — pacing, navigation depth, the order in which things are touched. A technically perfect request from a session that behaves wrongly is still refused.

What we doSession behaviour modelled on how a real client of that source actually moves, refitted whenever the scoring changes. That model is the asset; everything else is plumbing around it.

03 — RenderingRendering

Seat maps built in the browser

The seat map does not exist in the page. It is assembled client-side from a private interface, which is why most vendors reach for a browser and end up with a cost per refresh that makes minute-level freshness impossible.

What we doWe read the interface directly and keep a browser in reserve for the small share of cases where nothing else works. That single decision is most of the two-orders-of-magnitude gap in cost per refresh.

04 — DecayDecay

Value that expires in minutes

A price two minutes old is worthless for a hot event and perfectly fine for a show in eight months. Treating both the same way is either ruinously expensive or uselessly stale.

What we doEvents are ranked continuously by volatility and given a refresh budget accordingly, so freshness is spent where it changes a pricing decision.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Rank before you fetch

The first thing the pipeline does each cycle is decide what is worth looking at. Volatility, time to event, remaining inventory and the client's own exposure feed a ranking that allocates the next few minutes of capacity. Most systems in this space fetch first and think later, which is why their freshness collapses on exactly the days that matter.

02

One behaviour model, many identities

Identity rotation without a behaviour model is noise. We maintain a single model of how a legitimate client of this source behaves and instantiate it across identities, so every session is individually plausible rather than merely different from the last one.

03

Read the interface, not the page

Where the source exposes a private interface to its own front end, that is what we read. Browsers are expensive, fragile and slow, and reserving them for genuine exceptions is what keeps a minute-level refresh affordable at this catalogue size.

04

Reconcile every sweep

A second process samples the source independently — different entry points, different partitioning — and its result is compared against what the pipeline delivered. The difference is the coverage number. When it moves, we know within minutes rather than at the end of the quarter.

05

Deliver as a stream

Ninety-second data in a nightly file is ninety-second data nobody can use. Records land on a stream the moment they are reconciled, with a versioned schema, and the client's pricing engine consumes them directly.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

Price records a day155M
Event refreshes a day6.2M
Events on sale, tracked live500k
New events picked up a day5 800
Requests refused by the source0.4%

The arithmetic is the whole design. 500 000 events are on sale at any moment and 2.1 million distinct events pass through the catalogue in a year, because events go on sale and then happen. The refresh is tiered: 12 000 volatile events every three minutes over an eighteen-hour window is 4.3 million refreshes, 90 000 warm events twelve times a day adds 1.1 million, the remaining 400 000 are swept twice. That is 6.2 million refreshes a day — 72 a second, sustained. Each one returns the event's whole price map, about 25 rows, so 155 million price records a day, 4.7 billion a month, 57 billion a year and roughly 210 billion since 2019. Refreshing all 500 000 events every three minutes instead would be 180 million requests a day, and no client on earth could pay for it.

Headline

The four that are contractual

Collected since 2019
210B records
Freshness, hot tier
3 min
Coverage
99.2 %
Uptime, 12 mo
99.98 %

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Freshness on volatile events

Nightly batch, up to 14 h old

Three minutes

Catalogue coverage

62% of events, tail missing

99.2%, recounted daily

Decisions made on stale data

31% of repricing events

Under 2%

Engineering spent on collection

3 FTE, permanently firefighting

None on the client's side

The pipeline has run through two changes of bot vendor on the source, one full redesign of its front end, and a migration of the client's warehouse. It has not been rebuilt.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.