Send us the task
Sources 01 Defences 02 Scale 03 Architecture 04 Completeness 05 Reliability 06 Cases 07 Engagement 08 Questions 09 Send us the task

Atoms / Services / Complex data extraction

Discipline 01 — since 2014

The parsing everyone else calls impossible.

We build the extraction systems other teams quote as undoable: the most heavily defended sources in the industry, at millions of records a day, with coverage you can prove rather than hope for. This is the discipline we are number one in.

Twelve years, 1 000+ extraction systems and five billion delivered records. Have a task or a problem you need solved? Send it over, whatever its size — a great many of our clients arrived with something small and stayed for years. You get an answer within 24 hours.

Targetdefence stack / livePassing
LAYER 01 · EDGE / WAF Cloudflare · Akamai · Imperva LAYER 02 · BOT MANAGEMENT DataDome · HUMAN · Kasada LAYER 03 · FINGERPRINTING TLS · HTTP/2 · canvas · sensors LAYER 04 · CHALLENGES JS proofs · interstitials · queues LAYER 05 · LIMITS Rate caps · quotas · result caps ATOMS ACCESS Network routing Session strategy Fingerprint control Challenge pipeline Load shaping
Reach → Read → Reconcile → Deliver10M+ rows / day
1 000+

extraction systems built and shipped

5B+

records collected and delivered

10M+

records a day on a single pipeline

12+

years working only on hard sources

VerticalsTicketing · Real estate · Marketplaces · Travel · Mobility · Retail
Hardest classBehavioural bot management at national scale
EntrySend us any task or problem — an answer within 24 hours
Where we work01 / 10

The sources most teams don't quote.

We run production extraction against the most defended public platforms in their industries — the ones that hold an entire market's inventory and spend real money on keeping robots out.

TicketmasterTicketing

Event inventory, price tiers and availability on a platform with queue systems, aggressive behavioural scoring and per-event access rules.

Queue · behavioural · rate

ZillowReal estate

Listings, price history and status changes across a national portal with edge protection and result caps that hide most of the inventory from a naive crawl.

Edge WAF · result caps

StubHubSecondary market

Live seat-level pricing that moves by the minute, where the value of the data is entirely in how fresh it is.

Real time · high churn

AmazonMarketplace

Catalogue, offers, buy-box and stock at catalogue scale, with per-region and per-session variation in what is even shown.

Scale · personalisation

Booking.comTravel

Rates and availability that only exist as a function of a search — a combinatorial space, not a page you can list.

Query space · freshness

IdealistaReal estate EU

European portals with strict bot management, regional splits and short-lived listings that vanish before a weekly crawl sees them.

Bot management · churn

WalmartRetail

Store-level pricing and stock, where the answer differs per location and the interesting question is always local.

Geo matrix · volume

Your sourceUntested

The one another vendor returned as impossible. Two weeks, fixed fee, working proof.

Bring it to us

Named platforms are examples of the defence classes we run against in production. We collect publicly accessible data, we do not break authentication, and we shape load so the source is never degraded — the same rules on every engagement, written into the contract.

What stops everyone else02 / 10

Six defences, six answers.

A modern protected site is not one wall — it is six independent systems, each of which fails a scraper differently. Most projects die because one of them was never engineered for. Here is all six, and what we do about each.

01 — EdgeWAF

Managed bot rules at the CDN

Cloudflare, Akamai and Imperva score a request before the origin ever sees it — on network reputation, header order, protocol quirks and a hundred signals you don't get to read.

What we doCoherent clients rather than patched ones: transport, TLS and HTTP behaviour that is internally consistent, over routing chosen per source and per region.

02 — BehaviourScoring

Bot management that watches sessions

DataDome, HUMAN and Kasada don't judge a request, they judge a session: timing, navigation order, what a real user would have loaded on the way here.

What we doSession strategies modelled on real journeys, with pacing, warm-up and abandonment tuned per source instead of one global delay.

03 — IdentityFingerprint

Fingerprints across the whole stack

TLS handshakes, HTTP/2 frame settings, canvas and font enumeration, sensor entropy. A headless browser with a spoofed user-agent is identified instantly and quietly.

What we doFingerprint control end to end, so what the network layer claims and what the runtime does are the same story — and a drift monitor that tells us when the story stops working.

04 — ChallengeInterstitial

Proofs, interstitials and waiting rooms

JavaScript proof-of-work, invisible challenges, queue systems that hold a session for twenty minutes and expire it if it looks synthetic.

What we doA challenge pipeline that resolves in-session and keeps state through the wait, so throughput survives a queue instead of collapsing at it.

05 — StructureObfuscation

Markup engineered to be unparsable

Machine-generated class names that rotate on deploy, values assembled client-side, decoy nodes, content split across lazy fragments.

What we doWe work from the interfaces underneath the page wherever they exist, so a cosmetic redesign is a non-event rather than an outage.

06 — LimitsCaps

Ceilings that hide the catalogue

200 results per query over a catalogue of millions, pagination that stops at page 50, indexes that are deliberately partial. A naive crawl reports success and silently misses a third of the market.

What we doQuery-space decomposition — the catalogue is partitioned until every partition fits under the cap, then reconciled against control totals.

Scale03 / 10

Millions a day, entire sites, end to end.

Getting through a defence once is a demo. The engineering is in holding that throughput every day for years, at a cost per million rows that still makes the data worth having.

Operating envelope

What a single pipeline sustains

Full-site collection — the complete catalogue, not a sample — plus a delta pass fast enough that the picture is never more than a scheduled window old. Volume is a design input from the first day, not something discovered in month four when the pipeline stops keeping up.

Peak throughput
10M+ / day
Largest catalogue
180M+ items
Freshness floor
60 sec
Delivered to date
5B+ rows

Where the volume goes

Typical daily profile

Full catalogue sweep6.4M rows
Delta & price refresh2.9M rows
Real-time watchlist840k rows
Validation & reconciliationevery row
Rejected / quarantined0.3%

Illustrative profile of one production pipeline. Your envelope is fixed in the design phase and written into the SLA.

Architecture04 / 10

Built to survive the source fighting back.

Every parser we ship runs on the same contour. Blocks change per source; the property that does not change is that a failure anywhere is detected, isolated and replayed rather than silently delivered as missing rows.

Extraction contourrev. 07
Reference architecture of an Atoms extraction pipeline An orchestrator drives the schedule, queues and retries. Data moves from protected sources through an access layer, an extraction layer, validation and storage into delivery. An observability layer measures every stage and feeds corrections back into access and extraction. LAYER 00 Orchestration query-space partitioning · priorities · exponential back-off · idempotent jobs · replay SOURCE Protected Edge / WAFBot management ChallengesCaps & quotas LAYER 01 Access Routing & egressSession strategy Fingerprint controlChallenge pipeline adaptive throttling LAYER 02 Extract Interface-first parsingField mapping Rendering where neededSchema contracts drift detector LAYER 03 Verify & store Control totalsDeduplication Version historyQuarantine LAYER 04 Deliver API · webhooksS3 · warehouse Your databaseFeeds & files CROSS-CUTTING Observability per-stage metrics · block-rate tracking · coverage checks · alerting · on-call · automatic fallback route ROUTE CORRECTION SCHEMA REPAIR
InterfaceREST API
EventsWebhooks
Object storeS3
WarehouseBigQuery
WarehouseSnowflake
DirectYour database
Why we are number one at this05 / 10

Anyone can get a page. We prove we got all of them.

The difference between a scraping vendor and an extraction engineering team shows up months after launch — in whether the data is still arriving, and whether anyone can demonstrate that it is complete.

Dimension
Typical vendor
Atoms
Coverage
“It ran without errors,” with no way to know what the result caps hid.
Query space decomposed until nothing is capped, then reconciled against independent control totals and reported as a number.
The source changes
Silent breakage; you find out from a report that looks wrong two weeks later.
Drift detection on schema and volume fires before bad data lands, with contractual response windows.
Blocking
Rotate proxies harder until the cost per row makes the project pointless.
Block rate is a tracked metric with a target; access strategy is redesigned when it moves, not brute-forced.
Volume
Works in the pilot, falls over at production scale.
Throughput and cost per million rows are design inputs, fixed in the SLA before the build starts.
Data quality
Whatever the page happened to contain that day.
Schema contracts, range and type checks, dedupe and version history; failing batches quarantined, never published.
After go-live
The project closes, the team disperses, and the system quietly decays until something breaks badly enough to be noticed.
We run it. Hosting, monitoring, on-call, upstream changes and new features — month after month, by the people who built it.
01

Completeness is a measurement, not a claim

Every run is reconciled against control totals derived independently of the crawl — category counts, sitemap cardinality, known-item probes. Coverage ships as a number on the same dashboard we watch, so “did we get everything” is never a matter of opinion.

02

Whole sites, not sampled ones

We routinely take a source in full: hundreds of millions of items, every category, every region, every language variant. The partitioning strategy that makes that possible under result caps is the part that takes experience — it is why 1 000+ built systems matters more than any single clever trick.

03

Cost per million rows is the real constraint

On a defended source at scale, egress, compute and challenge handling are the recurring line item that decides whether the data is worth collecting. We design the execution model per source — interface-first where possible, rendered only where required — and report cost per million rows from day one.

04

We run what we build, for years

Tests, runbooks, infrastructure as code and a documented access strategy. Shipping is the start of it, not the end: hosting, monitoring, on-call, upstream changes, tuning and new features all sit inside one monthly arrangement, handled by the same engineers who designed the thing. A system we will still be running in year eight has to be engineered to be operated, not merely delivered — and most of our clients are on their third or fourth system with us.

Reliability06 / 10

What we commit to once the data flows.

A parser is a product with an uptime, not a script that was delivered once. These are the commitments in the support agreement; tiers and windows are set per project against how critical the feed is.

CommitmentTargetWhat it means in practice

Coverage

99 %+

Measured against independent control totals every run, not asserted. When coverage drops, you see it on the same dashboard we do, before it reaches a report.

Anomaly monitoring

24 / 7

Volume shifts, block-rate movement, schema drift, latency spikes and silent field loss are detected by us and raised to you — not discovered by your analyst.

Data quality checks

Every run

Completeness, duplicates, field types, ranges and cross-field consistency verified automatically; failing batches are quarantined rather than published.

Engineering response

From 4 h

Incident tiers with named windows. A source changing its defences or its markup is routine work covered by the retainer, not a change request.

Recovery

Defined RTO

Fallback routes, checkpoints and replay queues designed in advance, so an upstream change costs a window of latency — not a week of missing history.

Availability

99.9 %

Measured on delivered runs against schedule and reported monthly. You get the same dashboard we watch.

Selected work07 / 10

Three sources that had already failed.

Each of these came to us after another team returned it as impossible or handed over a pipeline that had stopped working. Clients are under NDA — the engineering detail we walk through against your own case.

01 — Ticketing2024

Live seat inventory across a queue-protected platform

Minute-level prices and availability for every event in nine countries, feeding a secondary-market pricing engine.

What was hardWaiting-room queues, behavioural session scoring, per-event access rules, and data whose value expires in under two minutes.

Events tracked
140k
Freshness
90 sec
Block rate
<0.4%
Downtime, 12 mo
0
02 — Real estate2025

A full national portal, not the first 50 pages

Complete daily inventory with price and status history across a country-scale listings portal behind managed bot protection.

What was hardA 1 000-result ceiling per query over 11M active listings, regional result variation, and a previous vendor's pipeline that had been silently missing 38% of the market for a year.

Coverage
99.4%
Listings / day
11M
Sweep window
5 h
Cost per 1k rows
−81%
03 — Marketplace2025

180M items, refreshed inside a day

Full catalogue, offers and stock for a global marketplace, delivered into the client's warehouse for a repricing engine.

What was hardCatalogue scale, per-region and per-session personalisation of what is even displayed, and a hard cost ceiling per million rows to keep the product viable.

Catalogue
180M
Peak / day
10M+
Full refresh
22 h
Field accuracy
99.8%
How we start08 / 10

We test your source before either side signs a build.

On a defended source, a quote written from a brief is a guess — the spread between the easy and the hard version of the same site is roughly tenfold. So we prove it on the real target first, for a fixed fee.

1Brief

The source and the decision

Which source, which fields, what volume, how fresh, and what the data is used for. Half the time the scope narrows and gets cheaper right here.

OutputTarget map & field spec

2Feasibility

Working proof, fixed fee

Two weeks against your actual target: access strategy, sample volume, real block rate, real cost per million rows — and the honest risks.

OutputRunning PoC + risk list

3Design

Schema, volume and SLA

Data model, delivery interface, throughput, freshness, coverage target and running cost — written into the contract, not into a slide.

OutputFixed scope & fixed price

4Build

Weekly working increments

Every week ships something that runs in your environment. Monitoring, runbooks and documentation are built alongside it, not bolted on at the end.

OutputA system running in production

5Operate

Monitoring and response

On-call, handling defence and markup changes on the source, tuning and extensions. Routine changes sit in the retainer; large ones are quoted first.

OutputA feed you stop thinking about

Engagement09 / 10

How we contract, and who we are wrong for.

Extraction prices by resistance, not by field count. So we price in two steps, and we say no early when the work is not ours.

Model

Two steps, both fixed

First task — one small piece of work, days rather than weeks, quoted before it starts and often all anyone needsFixed price
Feasibility check — two weeks on your real target; working proof, measured block rate and cost per million rowsFixed fee
Build — schema, coverage target, throughput, SLA and price fixed after the checkFixed price
Operate — hosting, monitoring, on-call, source-change handling, small extensionsMonthly

You start on whichever rung fits, and most clients start with one small first task. From the Operate rung onward it is a single monthly arrangement covering hosting, monitoring, on-call and continued development, so nobody on your side has to build an operations team around it. Infrastructure runs on ours or inside your own perimeter — whichever your security and finance people prefer.

Not our work

Say no early, honestly

  • One-off scrapes and single-page sites. We would be the expensive way to do that.
  • Bypassing authentication on accounts the client does not own, or anything that reads as unauthorised access.
  • Harvesting personal data without a lawful basis, and any load that would degrade the source.
  • Buying raw proxy capacity by the gigabyte. We take responsibility for delivered data, which requires owning the design.

If your source is real but not our kind of work, we will say so in the first reply and point you at someone better suited. That costs us nothing and saves you a month.

Legal & governance

Extraction your
legal team can sign off.

Projects are scoped around publicly accessible data, or access the client is authorised to hold. We document data purpose, minimise sensitive fields, and design access, retention and deletion controls into the pipeline rather than bolting them on.

GDPR and CCPA alignment is implemented through technical and organisational controls. Confirming the lawful basis and permitted use for a specific dataset and jurisdiction remains with the client’s legal team — we give them the documentation to do it.

01 — ScopePublic data, or access the client is authorised to hold
02 — LoadShaped so the source is never degraded
03 — PrivacyData minimisation and purpose limitation by design
04 — ContractNDA and DPA signed before technical detail is shared
Questions10 / 10

What gets asked about hard sources.

If yours isn’t here, put it in the form — we answer in writing, within 24 hours, and without a discovery call first. Most questions get a straight yes or no rather than a proposal.

Another team told us this source is impossible. Is it?
Usually it means the source is expensive rather than impossible, and the previous team priced it as if it were easy. That distinction is exactly what the feasibility check settles: two weeks on your actual target, a fixed fee, and a measured answer — block rate, achievable coverage, throughput and cost per million rows. Twelve years in, the question is almost never whether it can be done — it is what it costs to build and what it costs to keep running. You get both numbers in weeks rather than after a year.
Is this legal?
We work with publicly accessible data and with data the client is authorised to access. We do not break authentication, we do not collect personal data without a lawful basis, and we do not generate load that degrades a source. Constraints specific to your project and jurisdiction are assessed at the start and written into the agreement, so your legal team reviews a document rather than a promise.
What happens when the site changes its protection?
It is a question of when, not if, and it is designed for. Block rate and schema drift are monitored continuously, so a change is detected in minutes rather than in next month's report. Under a support agreement the response window is contractual and routine adaptation — a new challenge type, a markup rewrite, a changed rate policy — is covered by the retainer rather than quoted as new work.
How do you prove the data is complete?
Coverage is reconciled every run against control totals derived independently of the crawl — category counts, index cardinality, known-item probes — and reported as a number you can see. This is the single most common failure we inherit: a pipeline that ran without errors for a year while result caps quietly hid a third of the market.
Can you really do millions of records a day?
Yes, and the throughput number on its own is the easy part. What takes engineering is holding it at a cost per million rows that keeps the data worth collecting, while the source actively works against you. Both figures are measured during the feasibility check and fixed in the SLA before the build starts.
We have a parser that stopped working. Will you take it over?
Often, yes — a good share of our work starts as someone else's codebase. We audit it for a fixed fee and come back with one of three answers: stabilise it, keep the data model and rebuild the access layer, or replace. We will tell you which even when the honest answer is that the existing work is worth keeping and you need less from us than you thought.
Do you sell ready-made datasets or proxies?
No. We build extraction systems to your schema and your coverage target, and then we run them — hosting, monitoring, on-call, and the constant work of keeping up with a source that keeps changing. What you get is a feed your systems consume, not a shelf dataset. If what you actually need is a commodity dataset or raw proxy capacity, there are cheaper places to get it and we will say so.
How fast can we start?
A written feasibility assessment within three days of the brief, and the paid check itself usually starts within two weeks. Emergency work on a broken production feed is scheduled faster when we have the capacity — say so in the form and we will tell you honestly.

Bring us the hard source

Tell us which site
beat the last team.

Send the source, the fields and the volume you need. You get a straight answer: whether it is solvable, what specifically makes it hard, and roughly what it costs to run. No deck, no discovery call before there is anything to discover.

  • A substantive written reply within 24 hours, whatever the size
  • Feasibility assessment within three days of the brief
  • NDA signed before we go into technical detail
  • If it isn’t our kind of work, we say so immediately

By sending you agree to our privacy policy. Briefs are never shared and you will not be added to a mailing list.