Deep web research and RAG
Why web research is slow
Each new source needs its own scraper, and a manual research pass is hard to repeat. An agent that browses and summarizes gives you prose with no source trail. You cannot tell which claim came from which page, and a re-run gives a different answer.
Depth is the other gap. The facts you need are often spread across ten related pages, and a RAG index needs the whole site, PDFs included. The work is mapping a domain, fetching the relevant pages and reading them against one schema or one chunking scheme.
How SuperScraper fits
Define a schema, point the API at a URL or a domain and get structured output with the source URL on every result. For RAG, crawl the site and parse its documents into chunks.
- Structured fields without a parser
- /v1/extract with a schema returns typed fields under data. Without a schema it runs a free cascade: JSON-LD first, then Open Graph tags, then the page text. A language model runs only when you send a schema.
- A source URL on every result
- Every result carries the url it was read from, and a multi-URL call returns one result per URL. That list is your citation list.
- Many sources in one request
- /v1/extract accepts a urls array with one schema, so you pull the same fields from many sources at once. /v1/batch fetches pages as markdown when you want to read before you extract. Both go up to your plan’s cap.
- Documents for RAG
- /v1/parse turns a PDF or DOCX, by upload or URL, into markdown. Set format to rag to get chunks sized for embedding. There is no OCR, so a scanned PDF needs a text layer.
Endpoints
- POST
/v1/extractTyped fields in data against your schema, with the source URL. - POST
/v1/scrapeA page as clean markdown, to read before you write a schema. - POST
/v1/mapEvery URL on a domain, from its sitemap or homepage links. - POST
/v1/batchMany pages as markdown in one request, up to your plan’s cap. - POST
/v1/crawlA whole site as an async job, for large domains. - POST
/v1/parseA PDF or DOCX as markdown, text or RAG chunks. - POST
/v1/searchA query to ranked results. scrapeResults: true adds page content.
Extract facts with their source
Define the fields you want and pass the URL. The response returns the fields under data, with the source URL and model metadata. For many sources, send the same schema with a urls array.
Example responses use fictional data.
curl -X POST https://api.superscraper.dev/v1/extract \
-H "Authorization: Bearer $SS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/about",
"schema": {
"companyName": "string",
"founded": "string",
"headquarters": "string",
"products": "string[]"
}
}'
# Example response
{
"data": {
"companyName": "Example Corp",
"founded": "2012",
"headquarters": "Austin, TX",
"products": ["Platform A", "Platform B"]
},
"url": "https://example.com/about",
"metadata": { "model": "...", "latencyMs": 1900 }
}Questions
How does extraction work without a schema?
How do I get citations for extracted facts?
Can I research a whole domain?
Can I feed a RAG pipeline with it?
Which model runs schema extraction?
Start your first research workflow
Run any endpoint in the playground, or get a free key. 1,000 credits a month, no card.