Send us the source
Brief 01 What was hard 02 How it works 03 Numbers 04 Outcome 05 All cases → Send us the source

Atoms / Data extraction / Cases / Idealista

Real estate — running since 2023

Idealista — three countries, one property model.

A fund buying across southern Europe cannot compare a Madrid listing to a Lisbon one until both are the same shape. That means not just collecting three national markets, but resolving the same property where it appears twice and normalising three different ideas of what a square metre includes.

The source is named because it is public; the client is not. Every figure on this page is a measured production number the client agreed to publish. Have a source of your own? Send it over, whatever its size — you get an answer within 24 hours.

Case fileidealistaIn production
SourceIdealista ES / IT / PT
ClientEuropean property fund
Running since2023
DeliveryBigQuery, daily
Properties in the graph12M
Countries3
Reach → Read → Reconcile → DeliverRecounted daily
12B

property records to date

4.4B

a year

98.4%

measured coverage

12M

properties in the graph

SourceIdealista
VerticalReal estate
Running since2023 — without a rebuild
The brief01 / 05

What the client actually needed.

Not more rows. Every one of these projects started with somebody who already had data and could not use it for the decision in front of them.

The problem

Where it started

The client invests across three markets and had three data sources with three schemas, three definitions of floor area and no way to tell whether a property in one file was the same asset as a property in another. Every comparison in their investment committee pack was being made with a footnote apologising for the data.

Fixed on day one

What we committed to

All three countries in one schema with one area definitionContractual
Same asset resolved to one record where it appears more than onceRequired
Commercial and residential collected under the same modelBoth
Listing agent and price history retained for negotiation analysisFull

Every one of these is measured continuously and reported on the same dashboard the client watches. A commitment nobody measures is a sentence in a proposal.

What was hard02 / 05

Four things that beat the previous attempt.

None of these is solved by better headers or a bigger proxy pool. Each needs a different piece of engineering, and working out which one you are actually facing is most of the job.

01 — DefinitionsDefinitions

A square metre is not a square metre

Built area, usable area and plot area are reported differently in each market and sometimes inconsistently within one. A model that treats them as interchangeable produces valuations that are wrong by a third.

What we doArea fields are parsed into an explicit typed model with the source's own label retained, and derived comparables are computed only between like definitions.

02 — IdentityIdentity

The same asset, listed twice

Larger assets are listed by more than one agent, sometimes at different prices, and counting them twice distorts both supply figures and average price.

What we doEntity resolution on geometry, area, media fingerprints and description similarity, tuned against a hand-labelled set and re-measured weekly.

03 — MixMix

Commercial and residential together

The two have different fields, different search behaviour and different value drivers, and most collectors handle one and bolt the other on badly.

What we doOne property model with typed extensions per asset class, so the shared fields are genuinely shared and the differences are explicit rather than crammed into a notes column.

04 — GeographyGeography

Location precision varies

Some listings carry a precise point, some a neighbourhood polygon, some only a municipality — and a fund's model is extremely sensitive to which.

What we doLocation precision is a first-class field. A record carries what is known and how precisely, so the model can weight it rather than treating a municipality centroid as an address.

How it works03 / 05

Five decisions the pipeline is built on.

The architecture is not interesting; every extraction system has a queue, a fetcher and a parser. These are the decisions that made this one work where the last one did not.

01

Type the area fields

Built, usable and plot are separate typed fields with the source label preserved. Comparables are only ever computed between matching definitions, which is unglamorous and is the difference between a usable valuation and a wrong one.

02

Resolve on geometry and media

Duplicate detection leans on what is hard to change — location, area, photographs — rather than on text, which agents rewrite freely.

03

One model, typed extensions

Residential and commercial share a core and extend it explicitly. The alternative, two parallel pipelines, is how the two halves of a dataset drift apart over two years.

04

Carry precision, not just position

Every location record states how precisely it is known. Treating a municipality centroid as an address is the most common way property models quietly go wrong.

05

Recount per country

Coverage is measured separately for each country and each asset class. One weak country inside a healthy aggregate is exactly what per-country monitoring is for.

The numbers04 / 05

What it does on an ordinary day.

Production figures, not a benchmark run. Coverage is recounted daily against an independent sample of the live source rather than asserted, which is why the numbers are not round.

Daily profile

Where the volume goes

Property records a day12M
Requests a day3.2M
Active listings tracked1.6M
Duplicates collapsed144k
Location precision: exact point64%

Twelve million records a day is 360 million a month, 4.4 billion a year and 12 billion since 2023, from 3.2 million requests — 37 a second across three countries. Sixty-four per cent of listings carry an exact point and the rest do not; publishing that number, rather than silently geocoding the remainder to a centroid, is what lets the fund's model weight the difference.

Headline

The four that are contractual

Records to date
12B
Coverage
98.4 %
Dedup precision
97.6 %
Countries
3

These four sit in the support agreement. When one of them drifts outside its band, we are alerted within fifteen minutes and fixing it is routine work under the monthly arrangement, not a change request.

What changed05 / 05

Before, and after.

The columns are the client's own numbers from before the rebuild and the measured ones from production today. The left column is the part most vendors would rather not put on a page.

MetricBeforeToday

Cross-country comparison

With a footnote

Direct, one area definition

Duplicate assets

Counted twice

Resolved, sources retained

Commercial coverage

Separate, partial

In the same model

Location handling

Centroid where unknown

Precision as a field

The investment committee pack lost its data footnote. That was the actual deliverable, whatever the statement of work said.

Start here

Have a source of your own? Two lines are enough.

Send the source, the fields you need and roughly how often. You get a straight answer within 24 hours: whether it can be done, what makes it hard, what coverage is achievable and roughly what it costs to build and to run. Whatever the size.

Direct

NDA before technical detail, as always. If it isn't our kind of work, we say so in the first reply.