Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / Zillow

Real estate — running since 2021

Zillow — the whole market, not the first fifty pages.

The client had been buying a national listings feed for two years. It turned out to be missing roughly a third of the market — not randomly, but systematically, in exactly the segments their investment model cared about. We rebuilt it, and this time the coverage number is measured rather than claimed. Every new listing that appears anywhere in the country is in the warehouse the same day.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case filezillowIn production
SourceZillow + 40 regional portals
ClientProperty investment analytics
Running since2021
DeliverySnowflake + daily Parquet
Properties in the graph135M
Coverage SLA99% floor
Reach → Read → Reconcile → DeliverRecounted daily
26B

property records to date

7.3B

a year

99.1%

measured coverage

135M

properties in the graph

SourceZillow
VerticalReal estate
Running since2021 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

An investment analytics platform prices residential assets at national scale, so a gap in the inventory is not a data-quality annoyance — it is a hole in the model. The previous vendor delivered a large, plausible-looking file every night and reported success. Nobody had ever measured what share of the market that file actually represented, because measuring it is harder than collecting it.

Fixed on day one

What we committed to

Share of the live market present in the file each morning99% floor
One record per physical property, not one per portal listingContractual
Price and status history retained by us, since the source does not keep itFull
Whole-country refresh finished before the client's morning run5 h window

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — CeilingsCeilings

A hard result ceiling per query

The source will return a bounded number of results for any search, however the search is phrased. Over millions of active listings that ceiling is what quietly caps everybody's coverage — you collect a great many rows and have no idea what fraction of the whole they are.

What we doThe query space is partitioned — by geography, price band, property type, listing date — until every partition provably returns fewer results than the ceiling. Partitions that approach it are split again, automatically, as the market moves.

02 — VariationVariation

Results that differ by region

What the source returns depends on where it thinks the request comes from, which means a single vantage point produces a systematically regional view of a national market.

What we doVantage points are distributed to match the geography being swept, and each partition is verified from more than one so regional bias shows up as a discrepancy rather than as missing data.

03 — IdentityIdentity

The same property, four times

A property appears on the national portal, on two regional ones and on the agent's own site, with different identifiers, different photographs and different addresses spelled four ways. Counting those as four assets breaks a valuation model quietly.

What we doEntity resolution across sources using address normalisation, geometry, media fingerprints and listing history, tuned against a hand-labelled ground-truth set and re-measured every week.

04 — HistoryHistory

History the source does not keep

Price cuts, relistings and withdrawals are the most valuable signal in the dataset, and they are exactly what the source overwrites. If you were not watching yesterday, yesterday is gone.

What we doWe keep the history ourselves: every observed state is retained with its observation time, so price trajectories and time-on-market are reconstructed from our own record rather than trusted to the source.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Partition until it provably fits

Every partition carries a proof that it sits under the result ceiling, and the proof is re-checked on every sweep. When a city gets hot and its partition starts brushing the limit, it splits before anything is lost rather than after somebody notices a gap.

02

Sweep from where the market is

Vantage points follow the geography being collected, and every partition is verified from at least two so that regional variation surfaces as a discrepancy to investigate.

03

Resolve entities, then count

Deduplication happens before delivery, not in the client's warehouse. One physical property is one row, with the portals it appeared on as attributes, which is the only form in which the numbers can be trusted.

04

Keep your own history

Each observation is stored with its timestamp and never overwritten. Price history, status changes and time-on-market are derived from our own record, which is why the dataset gets more valuable the longer it runs.

05

Recount daily, publish the number

An independent process samples the live market each day and the delivered file is measured against it. The coverage figure in the client's dashboard is that measurement, not an estimate.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

Property records written a day20M
Requests a day7.4M
Properties refreshed a day4.5M
Active listings tracked1.4M
Price changes captured a day96k

Two collections run side by side. The live market — 1.4 million active listings — is swept completely inside five hours, 2.9 million requests at 160 a second. The standing graph of 135 million properties turns over on a rolling thirty-day cycle, 4.5 million properties a day, another 4.5 million requests spread across the remaining nineteen hours. Together that is 7.4 million requests a day — 86 a second — and 20 million written records, because one property read updates price, status, tax and attribute rows. Twenty million a day is 600 million a month, 7.3 billion a year and 26 billion since 2021.

Headline

The four that are contractual

Collected to date
26B records
Properties in the graph
135M
Coverage
99.1 %
History depth
5 yrs

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Share of the live market delivered

About 66%, unmeasured

99.1%, recounted daily

Records per physical property

Up to four, unresolved

One, with sources as attributes

Price history available

None — the source overwrites it

Five years and growing

Time to a full national refresh

Roughly three days

Five hours

The first month of the new pipeline surfaced whole segments the model had never seen. The valuation team's first reaction was that the numbers must be wrong; the recount is what settled it.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.