> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.eyelevel.ai/reference/api-reference/documents/crawl-website/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.eyelevel.ai/_mcp/server. # crawl_website POST https://api.groundx.ai/api/v1/ingest/documents/website Content-Type: application/json Upload the content of a publicly accessible website for ingestion into a GroundX bucket. This is done by following links within a specified URL, recursively, up to a specified depth or number of pages. Note1: This endpoint is currently not supported for on-prem deployments. Note2: The `source_url` must include the protocol, http:// or https://. [Supported Document Types and Ingest Capacities](https://docs.eyelevel.ai/documentation/fundamentals/document-types-and-ingest-capacities) Reference: https://docs.eyelevel.ai/reference/api-reference/documents/crawl-website ## Authentication - `X-API-Key` header (required) — API Key authentication via header ## Request ### Body (application/json) This endpoint expects a WebsiteCrawlRequest. - `websites` (list of WebsiteSource, required) - `callbackUrl` (string, optional) — The URL that will receive processing event updates. - `callbackData` (string, optional) — A string that is returned, along with processing event updates, to the callback URL. ## Response ### 200 Website successfully queued - `ingest` (IngestStatus, required) ## Errors ### 400 Bad Request Error Invalid source URL - `any` ### 401 Unauthorized Error Unauthorized to update bucket with given ID - `any` ## Types ### WebsiteSource - `bucketId` (integer, required) — The bucketId of the bucket which this website will be ingested into. - `sourceUrl` (string, required) — The URL from which the crawl is initiated. - `cap` (integer, optional) — The maximum number of pages to crawl - `depth` (integer, optional) — The maximum depth of linked pages to follow from the sourceUrl - `searchData` (WebsiteSourceSearchData, optional) — Custom metadata which can be used to influence GroundX's search functionality. This data can be used to further hone GroundX search. ### IngestStatus - `processId` (string, required) - `status` (enum, required) - Allowed values: `queued`, `training`, `generating`, `processing`, `error`, `complete`, `cancelled`, `active`, `inactive` - `id` (integer, optional) - `progress` (IngestStatusProgress, optional) - `statusMessage` (string, optional) ### WebsiteSourceSearchData Custom metadata which can be used to influence GroundX's search functionality. This data can be used to further hone GroundX search. ### IngestStatusProgress - `cancelled` (IngestStatusProgressCancelled, optional) - `complete` (IngestStatusProgressComplete, optional) - `errors` (IngestStatusProgressErrors, optional) - `processing` (IngestStatusProgressProcessing, optional) - `queued` (IngestStatusProgressQueued, optional) ### IngestStatusProgressCancelled - `documents` (list of DocumentDetail, optional) - `total` (integer, optional) ### IngestStatusProgressComplete - `documents` (list of DocumentDetail, optional) - `total` (integer, optional) ### IngestStatusProgressErrors - `documents` (list of DocumentDetail, optional) - `total` (integer, optional) ### IngestStatusProgressProcessing - `documents` (list of DocumentDetail, optional) - `total` (integer, optional) ### IngestStatusProgressQueued - `documents` (list of DocumentDetail, optional) - `total` (integer, optional) ### DocumentDetail - `documentId` (string, required) — Unique system generated ID for the document - `bucketId` (integer, optional) - `created` (string, optional) — Document creation time in RFC 3339 format, when available. - `deliveryStatus` (enum, optional) — Current-run required customer delivery confirmation. Failed confirmation does not prove nonreceipt; unknown means saved evidence is insufficient. - Allowed values: `not_configured`, `pending`, `succeeded`, `failed`, `unknown` - `extractionProvenance` (ExtractionProvenance, optional) — Workflow and runtime builds associated with the document's current extraction result. - `extractionReady` (boolean, optional, nullable) — Whether the current run's final extraction is published after required review and transformations. Null means saved evidence is insufficient. - `fileName` (string, optional) - `fileSize` (string, optional) — The file size of the file stored in GroundX - `fileType` (enum, optional) — The type of document (one of the currently supported file types) - Allowed values: `bmp`, `csv`, `docx`, `gif`, `heif`, `hwp`, `ico`, `jpg`, `json`, `pdf`, `png`, `pptx`, `svg`, `tiff`, `tsv`, `txt`, `xlsx`, `webp` - `filter` (DocumentDetailFilter, optional) — A dictionary of key-value pairs that can be used to pre-filter documents prior to a search. - `processId` (string, optional) — Unique system generated ID for the ingest request - `processLevel` (enum, optional) — The amount of processing of document chunks to perform. 'none' specifies the document to go through only text extraction and chunking. Documents uploaded with 'full' go through text extraction, chunking, and an agentic process that creates robust metadata for each chunk. (default 'full'). - Allowed values: `none`, `full` - `searchData` (DocumentDetailSearchData, optional) - `sourceUrl` (string, optional) — Source document URL - `hostedSourceUrl` (string, optional) — GroundX-hosted stored source URL, when available. Returned by document detail only; access still requires normal document authorization. - `status` (enum, optional) - Allowed values: `queued`, `training`, `generating`, `processing`, `error`, `complete`, `cancelled`, `active`, `inactive` - `statusMessage` (string, optional) - `textUrl` (string, optional) — Extracted text URL, if using the extract agent - `xrayUrl` (string, optional) — Document X-Ray results ### ExtractionProvenance Workflow and runtime builds associated with the document's current extraction result. - `artifactRevision` (string, optional) — Revision of the complete deployed workflow artifact, including prompts. - `runtimeBuildIds` (list of string, optional) — Unique extraction container build IDs observed while producing the current result. Each ID maps to one immutable image digest. - `schemaHash` (string, optional) — Hash of the workflow output schema and routing metadata. - `workflowId` (string, optional) — Unique system generated ID for the extraction workflow. - `workflowVersion` (string, optional) — Exact workflow version selected when extraction was dispatched. ### DocumentDetailFilter A dictionary of key-value pairs that can be used to pre-filter documents prior to a search. ### DocumentDetailSearchData ## Examples **Request** ```json { "websites": [ { "bucketId": 1234, "sourceUrl": "https://my.website.com", "cap": 10, "depth": 2, "searchData": { "key": "value" } } ] } ``` **Response** ```json { "ingest": { "processId": "uuid", "status": "queued" } } ``` **SDK Code** ```python import requests url = "https://api.groundx.ai/api/v1/ingest/documents/website" payload = { "websites": [ { "bucketId": 1234, "sourceUrl": "https://my.website.com", "cap": 10, "depth": 2, "searchData": { "key": "value" } } ] } headers = { "X-API-Key": "", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json()) ``` ```javascript const url = 'https://api.groundx.ai/api/v1/ingest/documents/website'; const options = { method: 'POST', headers: {'X-API-Key': '', 'Content-Type': 'application/json'}, body: '{"websites":[{"bucketId":1234,"sourceUrl":"https://my.website.com","cap":10,"depth":2,"searchData":{"key":"value"}}]}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); } ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.groundx.ai/api/v1/ingest/documents/website" payload := strings.NewReader("{\n \"websites\": [\n {\n \"bucketId\": 1234,\n \"sourceUrl\": \"https://my.website.com\",\n \"cap\": 10,\n \"depth\": 2,\n \"searchData\": {\n \"key\": \"value\"\n }\n }\n ]\n}") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("X-API-Key", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.groundx.ai/api/v1/ingest/documents/website") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["X-API-Key"] = '' request["Content-Type"] = 'application/json' request.body = "{\n \"websites\": [\n {\n \"bucketId\": 1234,\n \"sourceUrl\": \"https://my.website.com\",\n \"cap\": 10,\n \"depth\": 2,\n \"searchData\": {\n \"key\": \"value\"\n }\n }\n ]\n}" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.groundx.ai/api/v1/ingest/documents/website") .header("X-API-Key", "") .header("Content-Type", "application/json") .body("{\n \"websites\": [\n {\n \"bucketId\": 1234,\n \"sourceUrl\": \"https://my.website.com\",\n \"cap\": 10,\n \"depth\": 2,\n \"searchData\": {\n \"key\": \"value\"\n }\n }\n ]\n}") .asString(); ``` ```php request('POST', 'https://api.groundx.ai/api/v1/ingest/documents/website', [ 'body' => '{ "websites": [ { "bucketId": 1234, "sourceUrl": "https://my.website.com", "cap": 10, "depth": 2, "searchData": { "key": "value" } } ] }', 'headers' => [ 'Content-Type' => 'application/json', 'X-API-Key' => '', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.groundx.ai/api/v1/ingest/documents/website"); var request = new RestRequest(Method.POST); request.AddHeader("X-API-Key", ""); request.AddHeader("Content-Type", "application/json"); request.AddParameter("application/json", "{\n \"websites\": [\n {\n \"bucketId\": 1234,\n \"sourceUrl\": \"https://my.website.com\",\n \"cap\": 10,\n \"depth\": 2,\n \"searchData\": {\n \"key\": \"value\"\n }\n }\n ]\n}", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift import Foundation let headers = [ "X-API-Key": "", "Content-Type": "application/json" ] let parameters = ["websites": [ [ "bucketId": 1234, "sourceUrl": "https://my.website.com", "cap": 10, "depth": 2, "searchData": ["key": "value"] ] ]] as [String : Any] let postData = JSONSerialization.data(withJSONObject: parameters, options: []) let request = NSMutableURLRequest(url: NSURL(string: "https://api.groundx.ai/api/v1/ingest/documents/website")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```