POST
/web/crawlStart a website crawl
Submits a crawl job to extract content from pages on a website. Provide url as a URI and optionally set limit to cap the number of pages crawled; the default limit is 100. Use GET /web/crawl/{jobId} with the returned job ID to retrieve the crawl status and results.
- IdempotentThe SDK sends
Idempotency-Key, so a retried request is only applied once.
Website crawl configuration. Provide the required url and optionally set a maximum page limit.
urlstringrequired
URL of the website to crawl
limitnumberoptional
Maximum number of pages to crawl
200Returns the `jobId` for the crawl job that was started.
jobIdstringrequired
The ID of the job
400Returned when the crawl request is invalid.
errorstringrequired
Error code identifying the type of error
messagestringrequired
Human readable error message
detailsstringrequired
Detailed error description
documentationUrlstringoptional
URL to error documentation
429Returned when the request limit is exceeded.
errorstringrequired
Error code identifying the type of error
messagestringrequired
Human readable error message
detailsstringrequired
Detailed error description
documentationUrlstringoptional
URL to error documentation
500Returned when an internal error occurs.
errorstringrequired
Error code identifying the type of error
messagestringrequired
Human readable error message
detailsstringrequired
Detailed error description
documentationUrlstringoptional
URL to error documentation
Error handling
A 400 indicates an invalid request, a 429 indicates the request limit was exceeded, and a 500 indicates an internal error. The body must include url as a valid URI, and limit, when provided, must be between 1 and 5000.