Web data an agent can check.
SuperScraper gives a model a live, readable view of the web, and says how each page was read, so the model can decide what to trust before it acts.
Why a scrape should report on itself
Agents read the web at scale now, and most of the tools that feed them return a block of text with nothing about how it was obtained. The model cannot tell a clean read from a lucky one. A wrong price, a stale address or a parser that grabbed the wrong element goes downstream as if it were fact.
So every page scrape says how the page was fetched. When a site publishes structured data, we read it directly and say so, and we score how complete that data is. When a read comes back thin, the score is low, and your code can retry with a browser or hand the page to a model.
POST /v1/scrape → acme-plumbing.example
completeness78%
fetchplainreadjson-ldcompleteness0.78
What we hold ourselves to
- Checkable
- A page scrape says how the page was fetched. When the page has structured data, the result carries one completeness score for that data and says how it was read. A model can accept it, retry with a browser or ask a person.
- Fresh by default
- A scrape reads the live page unless you ask for a cached copy with maxAge. When a copy comes from cache, metadata.cached says so.
- Built for agents
- An agent creates its own free key with POST /v1/keys/provision, learns the API from a SKILL.md file and calls the REST API from its own loop.
- Respectful of the sites it reads
- Scrape and crawl respect robots.txt by default. A person can ask to have their details removed from company data through the takedown page.
See the readout on a real page
Paste a URL in the playground and look at the fetch method and completeness score. No key, no card.