Platform

Knowledge

Sources of text that agents search, such as web pages, handbooks and FAQs. A client writes a source as chunks, an officer uploads a document, or Platform crawls a URL on a schedule.

Access

  • Writers and searchers use a machine token with knowledge:write, knowledge:read or both. The org is the org of the token.
  • Search covers the caller's org and the public sources. A public source shows in the results of every org. Thus only publishers can write one. The superadmin marks publishers on the Superadmin page of the dashboard (PUT /api/superadmin/publishers/<org_id> with {"publisher": true}). Orgs in KNOWLEDGE_PUBLISHERS (org prefixes, comma-separated) are publishers too.
  • List, read and delete cover only the caller's org.
  • Officers upload documents, manage crawls, change the search settings and test search on the Knowledge page of the dashboard, with no token.

Routes

All routes are under /api/knowledge.

RouteScopeDoes
GET /sources?category=readThe org's sources and their current version
GET /sources/<key>readOne source
PUT /sources/<key>writeCreates or replaces a source. 201 if the content changed, 200 if not
DELETE /sources/<key>writeDeletes the source, its versions and its chunks
POST /searchreadRanked passages. The audit log does not record searches
PUT /crawls/<key>writeA crawled source: url, category, fetch_every_hours (1 to 720, default 24), title, public, enabled
POST /crawls/runwrite{"key": ..., "force": false}. Starts a crawl now. 202

The key is the writer's fixed name for the source: letters, digits and ._:/-, 255 characters or fewer.

{
  "category": "library",
  "title": "Library hours",
  "url": "https://example.edu/library/hours",
  "public": false,
  "content_hash": "optional; computed from the chunks if missing",
  "embedding_model": "required if the chunks have embeddings",
  "chunks": [
    {"ordinal": 0, "content": "...", "embedding": [1024 numbers], "level": 0, "parent_ordinal": 2}
  ]
}
  • A PUT replaces the whole source. If content_hash is the hash of the current version, only the title, URL, category and public flag change.
  • level 0 is page text. Levels 1 and higher are summaries of the rows below them. parent_ordinal points from a row to its summary. A source has up to 5000 chunks of up to 20000 characters.
  • All chunks have an embedding (1024 numbers), or none do. If none do and an embedder is set, Platform makes the embeddings. If there is no embedder, the source is text only.

The POST /search body has query, and optional category, top_k (1 to 50; the org's setting, default 8), window (0 to 5; the org's setting, default 0), and embedding with embedding_model to search with the caller's own vector. Each result has chunk_id, source_key, title, url, category, public, content, score and fetched_at. dense in the response shows if the vector search ran. The knowledge.search tool runs the same search.

Crawls

The knowledge.crawl_due job runs every 10 minutes. It crawls up to KNOWLEDGE_CRAWL_BATCH sources that are due, with KNOWLEDGE_CRAWL_GAP_SECONDS between fetches to the same host. Each crawl does these steps:

  1. Checks that the URL is http or https and resolves only to public addresses. It follows redirects one at a time and checks each one. Thus a source cannot point Platform at its own network.
  2. Reads robots.txt with KNOWLEDGE_USER_AGENT and skips pages that it does not allow.
  3. Gets the page through the org's Firecrawl (Integrations page, else FIRECRAWL_URL), else with a GET of up to 10 MB.
  4. Stops if the page hash did not change, unless force is set.
  5. Removes navigation, headers, footers, forms and scripts. Splits the text into chunks of the org's passage size with the page title on each, makes the embeddings and replaces the source's version. The old version and its chunks are deleted.
  6. Refuses the new text if it is less than half of the last version (when that was 500 characters or more), so a broken page cannot remove a good index. force accepts it.

A source's crawl field shows its schedule, last_attempt_at and last_error.

A page that starts to fail keeps its last chunks; the error shows on the source. Delete the source to remove them.

Uploads

POST /api/dashboard/<org>/knowledge/documents takes a multipart form: files (1 to 20 files, 10 MB each, 25 MB together), folder (default upload), category (default documents) and public. Platform reads .txt, .md, .markdown, .csv, .html, .htm, .pdf and .docx. Each file becomes a source with the key <folder>/<file name>, in lower case. The same name again replaces the source; if the file and the passage size did not change, nothing is written. A file that fails does not stop the others. The response lists each file with its key, passage count and error.

A PDF must have a text layer. Platform does not read scanned pages.

Run log

Each crawl and upload adds a row to knowledge_runs: the source key, crawl or upload, the time, the duration, if the index changed, the passage count and the error. Platform keeps the last 500 rows of each org. GET /api/dashboard/<org>/knowledge/runs?limit=&failed=1 reads them; the Activity page of the dashboard shows them in the Knowledge runs tab.

Org settings

Each org sets these on the Knowledge page (GET and PUT /api/dashboard/<org>/knowledge/settings). They are kept in the org config under knowledge. null returns a setting to its default.

SettingRangeDefaultDoes
chunk_chars100 to 4000KNOWLEDGE_CHUNK_CHARS, 300Passage size for crawls and uploads
chunk_overlap0 to half of chunk_chars0Text repeated from the end of the passage before
modehybrid, text, vectorhybridWhich searches run. Without an embedder, every mode uses text
top_k1 to 508Results when the caller sends no top_k
window0 to 50Neighbor rows when the caller sends no window
max_distance0.05 to 2KNOWLEDGE_MAX_DISTANCE, 0.6Vector results farther than this are dropped
rrf_k1 to 20060The constant of reciprocal rank fusion

A new passage size applies when a source is indexed again. POST /api/dashboard/<org>/knowledge/reindex starts the knowledge.reindex job, which crawls every crawled source of the org with force. Upload a document again to split it again.

Packs

A pack is a set of crawled pages and live queries that an org adds in one step, such as the ASU pack. The pack owns the org's sources whose keys start with <pack>/. The Knowledge page of the dashboard lists the packs (GET /api/dashboard/<org>/knowledge/packs) and syncs one (POST /api/dashboard/<org>/knowledge/packs/<name>/sync), which also starts the crawl job.

The dashboard groups sources by domain: the part of the key before the first /.

  1. Vector search: the chunks nearest to the query vector, of the same embedding model. It drops chunks farther than the org's max_distance (cosine distance). Mode text skips it.
  2. Text search: Postgres full text search (websearch_to_tsquery, ranked by ts_rank_cd). On SQLite, a count of shared words. Mode vector skips it when the vector search ran.
  3. Reciprocal rank fusion (k is the org's rrf_k) merges the lists.
  4. A row is dropped if its summary has a higher rank.
  5. With window, a page text result also has that number of rows on each side.

If the embedder fails during a search, the search uses text only. If it fails during a write, the write returns 502.

On Postgres with pgvector, the embedding column is vector(1024) with an HNSW index, and the text has a GIN index. If pgvector is not installed, the migration keeps the column as text and the vector search runs in Python. On SQLite both searches run in Python.

Settings

An org sets its own embeddings service and Firecrawl on the Integrations page of the dashboard (integrations). The variables below are the deployment defaults.

VariableDefaultDoes
EMBEDDINGS_URLnot setAn OpenAI-compatible base URL (.../v1). If neither the org nor this sets one, there is no embedder
EMBEDDINGS_MODELdefaultThe model name sent to the URL and kept on each version
EMBEDDINGS_API_KEYnot setThe bearer token for the URL
EMBEDDINGS_QUERY_PREFIXemptyText put before queries, for models that need it
KNOWLEDGE_MAX_DISTANCE0.6The default of the org setting max_distance
KNOWLEDGE_PUBLISHERSemptyOrg prefixes that can write public sources, in addition to the ones the superadmin marks
KNOWLEDGE_CHUNK_CHARS300The default of the org setting chunk_chars
KNOWLEDGE_CRAWL_BATCH20Sources crawled in each run of the job
KNOWLEDGE_CRAWL_GAP_SECONDS2The time between fetches to one host
KNOWLEDGE_USER_AGENTPlatformKnowledgeBot/1.0Sent with fetches and matched against robots.txt
FIRECRAWL_URL, FIRECRAWL_API_KEYnot setA Firecrawl for JavaScript pages

On this page