Last updated 2026-09-29
Where the data comes from
You point the API at a URL, a query or a domain. We read it at request time, and the response names how it was fetched and read.
Sources
| Source | How it is read | Reported in |
|---|---|---|
| The page you send | A scrape or extract call fetches the URL you pass, at request time. It reads from cache only when you send maxAge. We do not resell a stored copy of the web. | metadata.fetchMethod, metadata.cached |
| Structured data on that page | When a page publishes JSON-LD or Open Graph, we read those fields directly. No model runs and no credit is charged for that read. | listing._extraction_method, listing._completeness |
| Web search results | Search returns ranked results from a search index. With scrapeResults: true, each result page is scraped the same way as a direct scrape. | url on every result |
| Company websites | Resolve, brand and enrich read the company’s own site at request time, and reuse records kept from earlier lookups when they are recent. | _resolved_by, _freshness, _source_count |
How we handle it
- robots.txt respected by default
- Scrape and crawl honor a site’s robots.txt unless the request sets ignoreRobotsTxt: true, for content you are allowed to access.
- Live fetch by default
- A scrape reads the live page unless you send maxAge to accept a cached copy. A cached result says so in metadata.cached.
- The response says how
- metadata.fetchMethod says how a page was fetched. When structured data is found, _extraction_method says how it was read and _completeness scores the whole result from 0 to 1.
- Takedown on request
- A person can ask to have their details removed from company data. Removed records leave results right away and are permanently deleted after 30 days; a daily job completes the deletion, so it can take up to about a day longer. The takedown page says what to send.
Read the response fields
The docs show each field, which endpoints return it, and when it is present.