Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / Booking.com

Travel — running since 2022

Booking.com — a rate is a function, not a field.

A hotel revenue-management product is only as good as its view of the competition, and on an OTA the competitive rate is not a number sitting in a field. It is a function of date, length of stay, party size, market and currency — so collecting it means deciding, every day, which slice of a six-dimensional space is worth the money.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case filebookingIn production
SourceBooking.com
ClientHotel revenue-management platform
Running since2022
DeliveryBigQuery + hourly API
Properties1.6M
Markets24
Reach → Read → Reconcile → DeliverRecounted daily
130B

rate quotes since 2022

44B

a year

99.3%

rate accuracy

120M

quotes a day

SourceBooking.com
VerticalTravel
Running since2022 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client sells rate intelligence to hotels. Their previous approach collected a fixed grid — the next thirty days, two adults, one room — which is cheap, tidy and wrong the moment a hotel asks about a conference weekend or a seven-night family stay. Widening the grid naively multiplied the cost by forty and would have made the product unsellable.

Fixed on day one

What we committed to

Rate accuracy against a manual check of the live source99.5% floor
Markets covered at launch, with room to add without a rebuild38
Cost per million rows, agreed before the build and reported weeklyFixed
Refresh cadence for the dates that actually move revenueDaily

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — ExplosionExplosion

Six variables, one price

Check-in date, length of stay, occupancy, room type, market and currency all change the answer. The full space for a single property runs to millions of combinations, and almost all of them are worth nothing.

What we doWe model where the demand actually sits — per property, per season, per market — and sample the space against that model rather than gridding it. The client gets the combinations their hotels are asked about, not a uniform lattice.

02 — IdentityIdentity

What you are shown depends on who you seem to be

Prices, availability and even which properties appear vary with the session, the market and the device. Two observations taken a day apart are not comparable unless the observer was the same.

What we doSession identity is pinned per market and segment and held stable across time, so a change in the data means a change in the market rather than a change in who was asking.

03 — AvailabilityAvailability

Sold out is a data point

A missing rate can mean no availability, a closed restriction or a failed request, and treating those the same way corrupts every occupancy signal downstream.

What we doEach of the three is detected and recorded distinctly, with the evidence for the classification kept alongside the record so the client's model can trust the difference.

04 — CostCost

The number that decides whether the product exists

At this volume the recurring cost per million rows is not an operational detail — it is the difference between a product with a margin and one without.

What we doCost per million rows was a design input, fixed before the build. The sampling model, the request path and the caching of invariants were all chosen against that budget, and it is reported every week.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Model demand, then sample it

For each property we maintain a picture of which dates, stay lengths and party sizes actually get asked about, learned from the client's own query log and from seasonality. Capacity goes there. The uniform grid is the expensive way to be mostly wrong.

02

Pin the observer

One stable identity per market and segment, held across time. Comparability over months is worth more to a revenue model than any single day's breadth, and it is the first thing a rotating-identity approach destroys.

03

Separate the three kinds of nothing

No availability, a length-of-stay restriction and a failed request are recorded as three different things, with the evidence attached. Most feeds in this category collapse them into a null and quietly poison the occupancy signal.

04

Cache what does not move

Property attributes, room inventories and policies change on a scale of weeks; rates change hourly. Splitting the two and refreshing each on its own clock is where most of the cost reduction came from.

05

Report the cost weekly

Cost per million rows sits on the same dashboard as coverage and freshness. A pipeline whose unit economics drift is a pipeline that will be switched off in a year, so it is treated as a first-class metric.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

Rate quotes a day120M
Requests a day4.4M
Properties tracked1.6M
Cached, not re-fetched68%
Requests refused by the source0.9%

One request returns a block of dates, not a single price — about 27 quotes on average — so 4.4 million requests a day, 51 a second, produce 120 million quotes. That is 3.6 billion a month, 44 billion a year, 130 billion since 2022. Sixty-eight per cent of what the client needs is never fetched twice in a day at all, and that split is where the margin lives: pricing the same six-dimensional space by brute force would need thirty times the traffic.

Headline

The four that are contractual

Quotes since 2022
130B
Rate accuracy
99.3 %
Properties
1.6M
Cost / 1M quotes
−61 %

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Markets covered

4, expanding by hand

24, adding one is configuration

Rate combinations per property

A fixed 30-day grid

Sampled against real demand

Cost per million quotes

Baseline

61% lower

Accuracy against a manual check

Not measured

99.3%, sampled daily

The client's own pricing question — what does it cost us to answer one more market — went from a project to a line in a config file.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.