Skip to main content

Deep web research and RAG

Pull facts from public web pages into structured research, or turn whole sites and documents into chunks for a RAG pipeline. Every result carries the URL it was read from.

Why web research is slow

Each new source needs its own scraper, and a manual research pass is hard to repeat. An agent that browses and summarizes gives you prose with no source trail. You cannot tell which claim came from which page, and a re-run gives a different answer.

Depth is the other gap. The facts you need are often spread across ten related pages, and a RAG index needs the whole site, PDFs included. The work is mapping a domain, fetching the relevant pages and reading them against one schema or one chunking scheme.

How SuperScraper fits

Define a schema, point the API at a URL or a domain and get structured output with the source URL on every result. For RAG, crawl the site and parse its documents into chunks.

Structured fields without a parser
/v1/extract with a schema returns typed fields under data. Without a schema it runs a free cascade: JSON-LD first, then Open Graph tags, then the page text. A language model runs only when you send a schema.
A source URL on every result
Every result carries the url it was read from, and a multi-URL call returns one result per URL. That list is your citation list.
Many sources in one request
/v1/extract accepts a urls array with one schema, so you pull the same fields from many sources at once. /v1/batch fetches pages as markdown when you want to read before you extract. Both go up to your plan’s cap.
Documents for RAG
/v1/parse turns a PDF or DOCX, by upload or URL, into markdown. Set format to rag to get chunks sized for embedding. There is no OCR, so a scanned PDF needs a text layer.

Endpoints

  • POST/v1/extractTyped fields in data against your schema, with the source URL.
  • POST/v1/scrapeA page as clean markdown, to read before you write a schema.
  • POST/v1/mapEvery URL on a domain, from its sitemap or homepage links.
  • POST/v1/batchMany pages as markdown in one request, up to your plan’s cap.
  • POST/v1/crawlA whole site as an async job, for large domains.
  • POST/v1/parseA PDF or DOCX as markdown, text or RAG chunks.
  • POST/v1/searchA query to ranked results. scrapeResults: true adds page content.

Extract facts with their source

Define the fields you want and pass the URL. The response returns the fields under data, with the source URL and model metadata. For many sources, send the same schema with a urls array.

Example responses use fictional data.

curl -X POST https://api.superscraper.dev/v1/extract \
  -H "Authorization: Bearer $SS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/about",
    "schema": {
      "companyName": "string",
      "founded": "string",
      "headquarters": "string",
      "products": "string[]"
    }
  }'

# Example response
{
  "data": {
    "companyName": "Example Corp",
    "founded": "2012",
    "headquarters": "Austin, TX",
    "products": ["Platform A", "Platform B"]
  },
  "url": "https://example.com/about",
  "metadata": { "model": "...", "latencyMs": 1900 }
}

Questions

How does extraction work without a schema?
Without a schema, /v1/extract runs a free cascade: JSON-LD structured data first, then Open Graph tags, then patterns in the page text. The result carries data._completeness, a 0 to 1 score for how complete the structured data is, and data._extraction_method, which names the layer that produced it. A language model runs only when you send a schema, and that response carries data and model metadata instead of a completeness score.
How do I get citations for extracted facts?
Every /v1/extract and /v1/scrape response includes the url the data came from. With a urls array on /v1/extract, or with /v1/batch, each result carries its own url. Use those as the citation list.
Can I research a whole domain?
Yes. Use /v1/map to list the URLs on a domain, then send them to /v1/extract with a schema or fetch them with /v1/batch. For large sites, /v1/crawl runs as an async job: it returns a jobId you poll with GET /v1/crawl/:id.
Can I feed a RAG pipeline with it?
Yes. Crawl the site to get its pages as markdown, and send its PDF and DOCX files to /v1/parse with format set to rag to get chunks. Embed the chunks and store them in your own vector store.
Which model runs schema extraction?
A router picks the model from the task complexity and your plan. Set complexity to "high" for harder pages to get a larger model. metadata.model and metadata.provider in the response say which model ran.

Start your first research workflow

Run any endpoint in the playground, or get a free key. 1,000 credits a month, no card.