> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.eyelevel.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.eyelevel.ai/_mcp/server.

# crawl_website

POST https://api.groundx.ai/api/v1/ingest/documents/website
Content-Type: application/json

Upload the content of a publicly accessible website for ingestion into a GroundX bucket. This is done by following links within a specified URL, recursively, up to a specified depth or number of pages.

Note1: This endpoint is currently not supported for on-prem deployments. 
Note2: The `source_url` must include the protocol, http:// or https://.

[Supported Document Types and Ingest Capacities](https://docs.eyelevel.ai/documentation/fundamentals/document-types-and-ingest-capacities)


Reference: https://docs.eyelevel.ai/reference/api-reference/documents/crawl-website

## Authentication

- `X-API-Key` header (required) — API Key authentication via header

## Request

### Body (application/json)

This endpoint expects a WebsiteCrawlRequest.

- `websites` (list of WebsiteSource, required)
- `callbackUrl` (string, optional) — The URL that will receive processing event updates.
- `callbackData` (string, optional) — A string that is returned, along with processing event updates, to the callback URL.

## Response

### 200

Website successfully queued

- `ingest` (IngestStatus, required)

## Errors

### 400 Bad Request Error

Invalid source URL

- `any`

### 401 Unauthorized Error

Unauthorized to update bucket with given ID

- `any`

## Types

### WebsiteSource

- `bucketId` (integer, required) — The bucketId of the bucket which this website will be ingested into.
- `sourceUrl` (string, required) — The URL from which the crawl is initiated.
- `cap` (integer, optional) — The maximum number of pages to crawl
- `depth` (integer, optional) — The maximum depth of linked pages to follow from the sourceUrl
- `searchData` (WebsiteSourceSearchData, optional) — Custom metadata which can be used to influence GroundX's search functionality. This data can be used to further hone GroundX search.

### IngestStatus

- `processId` (string, required)
- `status` (enum, required)
  - Allowed values: `queued`, `training`, `generating`, `processing`, `error`, `complete`, `cancelled`, `active`, `inactive`
- `id` (integer, optional)
- `progress` (IngestStatusProgress, optional)
- `statusMessage` (string, optional)

### WebsiteSourceSearchData

Custom metadata which can be used to influence GroundX's search functionality. This data can be used to further hone GroundX search.

### IngestStatusProgress

- `cancelled` (IngestStatusProgressCancelled, optional)
- `complete` (IngestStatusProgressComplete, optional)
- `errors` (IngestStatusProgressErrors, optional)
- `processing` (IngestStatusProgressProcessing, optional)
- `queued` (IngestStatusProgressQueued, optional)

### IngestStatusProgressCancelled

- `documents` (list of DocumentDetail, optional)
- `total` (integer, optional)

### IngestStatusProgressComplete

- `documents` (list of DocumentDetail, optional)
- `total` (integer, optional)

### IngestStatusProgressErrors

- `documents` (list of DocumentDetail, optional)
- `total` (integer, optional)

### IngestStatusProgressProcessing

- `documents` (list of DocumentDetail, optional)
- `total` (integer, optional)

### IngestStatusProgressQueued

- `documents` (list of DocumentDetail, optional)
- `total` (integer, optional)

### DocumentDetail

- `documentId` (string, required) — Unique system generated ID for the document
- `bucketId` (integer, optional)
- `created` (string, optional) — Document creation time in RFC 3339 format, when available.
- `deliveryStatus` (enum, optional) — Current-run required customer delivery confirmation. Failed confirmation does not prove nonreceipt; unknown means saved evidence is insufficient.
  - Allowed values: `not_configured`, `pending`, `succeeded`, `failed`, `unknown`
- `extractionProvenance` (ExtractionProvenance, optional) — Workflow and runtime builds associated with the document's current extraction result.
- `extractionReady` (boolean, optional, nullable) — Whether the current run's final extraction is published after required review and transformations. Null means saved evidence is insufficient.
- `fileName` (string, optional)
- `fileSize` (string, optional) — The file size of the file stored in GroundX
- `fileType` (enum, optional) — The type of document (one of the currently supported file types)
  - Allowed values: `bmp`, `csv`, `docx`, `gif`, `heif`, `hwp`, `ico`, `jpg`, `json`, `pdf`, `png`, `pptx`, `svg`, `tiff`, `tsv`, `txt`, `xlsx`, `webp`
- `filter` (DocumentDetailFilter, optional) — A dictionary of key-value pairs that can be used to pre-filter documents prior to a search.
- `processId` (string, optional) — Unique system generated ID for the ingest request
- `processLevel` (enum, optional) — The amount of processing of document chunks to perform. 'none' specifies the document to go through only text extraction and chunking. Documents uploaded with 'full' go through text extraction, chunking, and an agentic process that creates robust metadata for each chunk. (default 'full').
  - Allowed values: `none`, `full`
- `searchData` (DocumentDetailSearchData, optional)
- `sourceUrl` (string, optional) — Source document URL
- `hostedSourceUrl` (string, optional) — GroundX-hosted stored source URL, when available. Returned by document detail only; access still requires normal document authorization.
- `status` (enum, optional)
  - Allowed values: `queued`, `training`, `generating`, `processing`, `error`, `complete`, `cancelled`, `active`, `inactive`
- `statusMessage` (string, optional)
- `textUrl` (string, optional) — Extracted text URL, if using the extract agent
- `xrayUrl` (string, optional) — Document X-Ray results

### ExtractionProvenance

Workflow and runtime builds associated with the document's current extraction result.

- `artifactRevision` (string, optional) — Revision of the complete deployed workflow artifact, including prompts.
- `runtimeBuildIds` (list of string, optional) — Unique extraction container build IDs observed while producing the current result. Each ID maps to one immutable image digest.
- `schemaHash` (string, optional) — Hash of the workflow output schema and routing metadata.
- `workflowId` (string, optional) — Unique system generated ID for the extraction workflow.
- `workflowVersion` (string, optional) — Exact workflow version selected when extraction was dispatched.

### DocumentDetailFilter

A dictionary of key-value pairs that can be used to pre-filter documents prior to a search.

### DocumentDetailSearchData

## Examples

**Request**

```json
{
  "websites": [
    {
      "bucketId": 1234,
      "sourceUrl": "https://my.website.com",
      "cap": 10,
      "depth": 2,
      "searchData": {
        "key": "value"
      }
    }
  ]
}
```

**Response**

```json
{
  "ingest": {
    "processId": "uuid",
    "status": "queued"
  }
}
```

**SDK Code**

```python
import requests

url = "https://api.groundx.ai/api/v1/ingest/documents/website"

payload = { "websites": [
        {
            "bucketId": 1234,
            "sourceUrl": "https://my.website.com",
            "cap": 10,
            "depth": 2,
            "searchData": { "key": "value" }
        }
    ] }
headers = {
    "X-API-Key": "<apiKey>",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
```

```javascript
const url = 'https://api.groundx.ai/api/v1/ingest/documents/website';
const options = {
  method: 'POST',
  headers: {'X-API-Key': '<apiKey>', 'Content-Type': 'application/json'},
  body: '{"websites":[{"bucketId":1234,"sourceUrl":"https://my.website.com","cap":10,"depth":2,"searchData":{"key":"value"}}]}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go
package main

import (
	"fmt"
	"strings"
	"net/http"
	"io"
)

func main() {

	url := "https://api.groundx.ai/api/v1/ingest/documents/website"

	payload := strings.NewReader("{\n  \"websites\": [\n    {\n      \"bucketId\": 1234,\n      \"sourceUrl\": \"https://my.website.com\",\n      \"cap\": 10,\n      \"depth\": 2,\n      \"searchData\": {\n        \"key\": \"value\"\n      }\n    }\n  ]\n}")

	req, _ := http.NewRequest("POST", url, payload)

	req.Header.Add("X-API-Key", "<apiKey>")
	req.Header.Add("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("https://api.groundx.ai/api/v1/ingest/documents/website")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["X-API-Key"] = '<apiKey>'
request["Content-Type"] = 'application/json'
request.body = "{\n  \"websites\": [\n    {\n      \"bucketId\": 1234,\n      \"sourceUrl\": \"https://my.website.com\",\n      \"cap\": 10,\n      \"depth\": 2,\n      \"searchData\": {\n        \"key\": \"value\"\n      }\n    }\n  ]\n}"

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.groundx.ai/api/v1/ingest/documents/website")
  .header("X-API-Key", "<apiKey>")
  .header("Content-Type", "application/json")
  .body("{\n  \"websites\": [\n    {\n      \"bucketId\": 1234,\n      \"sourceUrl\": \"https://my.website.com\",\n      \"cap\": 10,\n      \"depth\": 2,\n      \"searchData\": {\n        \"key\": \"value\"\n      }\n    }\n  ]\n}")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.groundx.ai/api/v1/ingest/documents/website', [
  'body' => '{
  "websites": [
    {
      "bucketId": 1234,
      "sourceUrl": "https://my.website.com",
      "cap": 10,
      "depth": 2,
      "searchData": {
        "key": "value"
      }
    }
  ]
}',
  'headers' => [
    'Content-Type' => 'application/json',
    'X-API-Key' => '<apiKey>',
  ],
]);

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("https://api.groundx.ai/api/v1/ingest/documents/website");
var request = new RestRequest(Method.POST);
request.AddHeader("X-API-Key", "<apiKey>");
request.AddHeader("Content-Type", "application/json");
request.AddParameter("application/json", "{\n  \"websites\": [\n    {\n      \"bucketId\": 1234,\n      \"sourceUrl\": \"https://my.website.com\",\n      \"cap\": 10,\n      \"depth\": 2,\n      \"searchData\": {\n        \"key\": \"value\"\n      }\n    }\n  ]\n}", ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let headers = [
  "X-API-Key": "<apiKey>",
  "Content-Type": "application/json"
]
let parameters = ["websites": [
    [
      "bucketId": 1234,
      "sourceUrl": "https://my.website.com",
      "cap": 10,
      "depth": 2,
      "searchData": ["key": "value"]
    ]
  ]] as [String : Any]

let postData = JSONSerialization.data(withJSONObject: parameters, options: [])

let request = NSMutableURLRequest(url: NSURL(string: "https://api.groundx.ai/api/v1/ingest/documents/website")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers
request.httpBody = postData as Data

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```