Skip to main content

Web data an agent can check.

SuperScraper gives a model a live, readable view of the web, and says how each page was read, so the model can decide what to trust before it acts.

Why a scrape should report on itself

Agents read the web at scale now, and most of the tools that feed them return a block of text with nothing about how it was obtained. The model cannot tell a clean read from a lucky one. A wrong price, a stale address or a parser that grabbed the wrong element goes downstream as if it were fact.

So every page scrape says how the page was fetched. When a site publishes structured data, we read it directly and say so, and we score how complete that data is. When a read comes back thin, the score is low, and your code can retry with a browser or hand the page to a model.

POST /v1/scrape → acme-plumbing.example

completeness78%
fetchplainreadjson-ldcompleteness0.78
Example readout for a fictional business. The values sit in metadata.fetchMethod, listing._extraction_method and listing._completeness of a scrape response.

What we hold ourselves to

Checkable
A page scrape says how the page was fetched. When the page has structured data, the result carries one completeness score for that data and says how it was read. A model can accept it, retry with a browser or ask a person.
Fresh by default
A scrape reads the live page unless you ask for a cached copy with maxAge. When a copy comes from cache, metadata.cached says so.
Built for agents
An agent creates its own free key with POST /v1/keys/provision, learns the API from a SKILL.md file and calls the REST API from its own loop.
Respectful of the sites it reads
Scrape and crawl respect robots.txt by default. A person can ask to have their details removed from company data through the takedown page.

See the readout on a real page

Paste a URL in the playground and look at the fetch method and completeness score. No key, no card.