WebRobot maintains a curated catalogue of ready-made extraction recipes — one per source — so your agents build data pipelines without rediscovering how each site works. The catalogue refreshes and repairs itself on a schedule.
source→recipe→pipeline→data— zero live inference when the recipe is healthyFor each source the catalogue has already resolved how to acquire it — so the agent doesn't guess.
Simple HTTP vs a real browser, its anti-bot posture, and whether a residential proxy / geo is needed.
The repeating element (product card, thread row) and its fields — title, price, link, author…
The site's own search box, so a query lands on a real results page before extraction.
Next-link, load-more or URL pattern — to go beyond page one.
The link to each item's detail page, followed with a join when per-item detail is needed.
Every recipe is proven to extract real rows before it enters the catalogue.
A scheduled loop keeps the catalogue honest — it checks first, and only rebuilds what actually broke.
Each source's saved recipe runs against the live page and is judged coherently — not just “did it return rows”, but: is the row count plausible versus the recipe's baseline, and do the key fields carry sane values (a URL that is a link, a price that looks like a price)?
A page that is empty or challenged (anti-bot, geo, proxy, HTTP error) is an access problem, not a broken recipe — it is left untouched and flagged, never re-inferred by mistake.
When a site genuinely changed its markup, that one source is re-set-up in a single browser pass — fetch class, item list, search, pagination and detail — and the repaired recipe is written back.
Everything that still validates stays exactly as it is. Cost scales with how much broke, not with catalogue size.
navigate→search→wait to render→infer item list→fields→pagination→detailSetup drives a real browser and combines deterministic detection with LLM inference for precision. For shops, the search is performed first so the item list is learned on a real results page, not the homepage. JavaScript-heavy pages are given time to hydrate before anything is read, and every search / pagination control is kept only once it is verified by actually driving it in the browser.
A cataloged source assembles a working pipeline with zero live inference — no per-run cost to rediscover selectors.
Recipes are pre-validated and watched; when a site drifts, the platform notices and repairs it — not your pipeline at 3am.
The agent uses the recipe directly on a healthy source, and re-infers live only on a genuine miss.
The catalogue is a private platform asset, read per tenant with your own credential — never a shared key, and never public. The pre-inferred recipes are part of a WebRobot subscription; the live demo shows the same engine inferring a source from scratch.