# Ingestion API

> REST endpoints for pushing content into a Reaktly knowledge base.

URL: https://docs.reaktly.com/docs/api-reference/ingestion

## Base URL

```
https://api.reaktly.com
```

Every request needs an API key with the `iq:import` scope (see [Authentication](/docs/api-reference/authentication)):

```bash
curl https://api.reaktly.com/ingest \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_API_KEY" \
  -d '{ ... }'
```

Ingestion is asynchronous: an accepted item is queued, processed (normalized, chunked, embedded) and then searchable by the AI. `jobId`s are derived from `integrationId` + `externalId`, so re-submitting the same pair does not create duplicate work.

## POST /ingest — single item

### Request body

| Field | Type | Required | Description |
|---|---|---|---|
| `integrationId` | string | Yes | Your integration's ID (the connector this data belongs to) |
| `knowledgeBaseId` | string | No | Target knowledge base. Omit to use the tenant's default |
| `sourceType` | string | Yes | Content kind — see [Source types](/docs/integrations/source-types) |
| `externalId` | string | Yes | Stable ID from your system; the deduplication key within the integration |
| `sourceOrigin` | string | No | Origin tag (e.g. `wordpress-rest`). Scopes [cleanup](#post-ingestcleanup--remove-orphaned-items) to this connector |
| `data.title` | string | Yes | Item title |
| `data.content` | string | Yes | The text the AI answers from — this is what gets embedded |
| `data.url` | string | No | Canonical URL of the original item, used in citations |
| `data.metadata` | object | No | Original raw data (price, sku, author, …) |
| `data.sourceData` | object | No | Structured source data for richer answers |
| `data.hash` | string | No | Content hash — unchanged hashes skip re-processing ([Change detection](/docs/integrations/change-detection)) |
| `options.forceRefresh` | boolean | No | Re-process even if the hash is unchanged |
| `options.priority` | `"high"` \| `"low"` | No | Queue priority. Use `high` for user-triggered imports, `low` for background syncs |
| `options.metadataOnly` | boolean | No | Update `url`, `metadata` and `sourceData` without re-embedding — for price or URL changes |

### Example

```json
{
  "integrationId": "int_abc123",
  "knowledgeBaseId": "kb_xyz789",
  "sourceType": "ARTICLE",
  "externalId": "blog-post-42",
  "sourceOrigin": "my-cms",
  "data": {
    "title": "How to use our product",
    "content": "This guide walks you through...",
    "url": "https://your-site.com/blog/how-to-use",
    "hash": "a1b2c3d4e5f6"
  },
  "options": {
    "priority": "high"
  }
}
```

### Response `200`

```json
{
  "status": "queued",
  "jobId": "int_abc123-blog-post-42"
}
```

## POST /ingest/bulk — batch import

Optimised for initial loads. Each item has the shape above; the tenant is taken from the API key.

```json
{
  "items": [
    { "integrationId": "int_abc123", "sourceType": "PRODUCT", "externalId": "sku-1", "data": { "title": "…", "content": "…" } },
    { "integrationId": "int_abc123", "sourceType": "PRODUCT", "externalId": "sku-2", "data": { "title": "…", "content": "…" } }
  ]
}
```

### Response `200`

```json
{
  "status": "queued",
  "jobCount": 50,
  "estimatedTime": "5s",
  "jobIds": ["int_abc123-sku-1", "int_abc123-sku-2"]
}
```

> **Note:**
> `estimatedTime` is the queue's own estimate (10 items per second). It describes queueing, not embedding time.

## POST /ingest/cleanup — remove orphaned items

After a full sync, tell Reaktly which `externalId`s still exist. Sources with the same `sourceOrigin` that are **not** in the keep list are deleted — so items removed from your CMS disappear from the knowledge base too.

```json
{
  "knowledgeBaseId": "kb_xyz789",
  "sourceOrigin": "my-cms",
  "keepExternalIds": ["blog-post-42", "blog-post-43"]
}
```

### Response `200`

```json
{
  "deleted": 3,
  "orphanedExternalIds": ["blog-post-17", "blog-post-18", "blog-post-19"]
}
```

Cleaning up is scoped: only sources whose `sourceOrigin` matches are considered. An empty `keepExternalIds` list deletes every source from that origin — use it with care.

## Errors

See [Error handling](/docs/integrations/error-handling) for status codes, retry and backoff.

## Where to go next

- [Data ingestion guide](/docs/integrations/ingestion-api) — worked examples in TypeScript, Python and cURL
- [Source types](/docs/integrations/source-types) — how to classify what you send
- [Change detection](/docs/integrations/change-detection) — hashing to skip unchanged content