What it does
- Reads every markdown file in the hub and parses YAML frontmatter against the hub schema.
- Chunks each body by heading.
- Embeds chunks via OpenAI (
text-embedding-3-smallat 768 dimensions, the default since 2026-09-03), Ollama (nomic-embed-text), or a deterministic local-hash fallback for CI. - Upserts chunks into Qdrant with payload indices on
kh_id,scope,scope_prefixes,sensitivity,source_type,team, andtags. - Maintains a Postgres manifest so unchanged files are skipped on subsequent runs.
- Sweeps orphans — IDs in the manifest but no longer in the hub get deleted from both Qdrant and the manifest.
Embedding providers
- OpenAI (default since 2026-09-03):
text-embedding-3-smallatdimensions=768, requested server-side so the vector arrives unit-length — required, because the collections are Cosine and a client-side slice would need re-normalizing. - Ollama:
nomic-embed-texton the GPU node. Still supported, and no longer used: it made every KB query wait on a node that scales to zero. - Hash: deterministic local fallback for CI smoke tests; never used in production.
config/model_catalog.yaml
(kind: embedding, dimensions: 768), credential and proxy resolution lives in
worker/worker/llm/embeddings.py, and callers that cannot reach a provider
themselves ask telemetry-api’s POST /v1/kb/embed rather than growing a second
implementation. Changing the model or the width is a re-embed, not an edit.
Two things the hosted path needs that the local one did not, both learned by
having them missing: the credential is in Secrets Manager (OPENAI_API_KEY is
present in the environment and empty, so a check that stopped there reported
“not configured”), and OpenAI geo-blocks ap-east-1 — every call leaves
through the proxy named by the provider’s proxy_env.
Batched. EmbeddingProvider.embed_many sends up to
OPENAI_EMBED_BATCH_SIZE (128) inputs per request, which is what makes a
~30,000-chunk backfill a few hundred round trips instead of 30,000. Order is
restored from each item’s index rather than assumed: vectors are zipped onto
chunks positionally, so a reordered response would attach every vector to the
wrong document.
CLI
Configuration
The manifest write is skipped under
--no-manifest (first runs, debug, or when Postgres is unavailable) and under --dry-run.