What it does
- Pulls jobs from Redis.
- Runs the context engine before routing — pulls knowledge-hub augments and RAG-retrieved snippets into the prompt.
- Routes paid jobs through
ScenarioOrchestrator; see Scenarios runtime for the engine and Scenario execution for the per-level walkthrough. - Tracks escalation state and per-conflict retry counts.
- Emits typed
AgentEventJSON to a Redis list and pubsub channel — scenario lifecycle (scenario_*,level_*,verifier_*,gate_decision), routing (routing_decisionwith the per-level band,model_fallbackwhen the band walks), approvals (approval_requested,ask_user), budget (budget_*), and artifacts (file_created,final,vercel_deployed). Payload contracts live in HTTP API & event stream. - Persists status, cost, and audit trail.
- Parks every human gate as a durable
approvalsrow — in-graph scope/clarify gates,ask_usertool calls, and the pre-graph confirmation gates alike — so a decision can land from the Action Dock, the/approvalspage, or Slack, not only from the live stream. - Provisions per-project workspaces (clone, scaffold, trigger Vercel deploy).
Retrieval collections
The worker reads three Qdrant collections and owns two of them. What this enables. User-authored rules and facts are retrievable at all, and a retrieval that fails is now distinguishable from one that found nothing. Previouslycerebrum_user_rules and cerebrum_facts came into existence only as a side effect of the
first successful rule/fact write; until that landed, every query returned a 404 that the
read path logged as a warning and converted into an empty list. The context node reported
success either way.
Who provisions what.
The preflight is
ensure_runtime_collections() in worker/worker/rag/kb_writer.py, called
from worker/worker/main.py at boot. It is idempotent, non-fatal, and self-healing: a
worker that cannot reach Qdrant still boots, and a rebuilt Qdrant is re-provisioned on the
next rollout.
The width is configuration, not whichever vector arrives first. Qdrant cannot resize a
collection, so the width it is born with is permanent. Collections are created at
EmbeddingProvider.vector_size — EMBED_DIMENSIONS, else the model catalog — rather than
len(vector). A stored width that disagrees with the configured one is logged as an ERROR
and left alone; recreating it is data loss and an operator’s call. EMBEDDING_PROVIDER=hash
is a CI-only mode whose vectors are 384-d, so the preflight refuses to provision against a
non-local Qdrant.
Failures are visible but never fatal. _retrieve_from_collection classifies a missing
collection separately from a query failure, logs both at ERROR, and still degrades to an
empty result rather than killing the job. AugmentedPromptPayload carries
retrieval_degraded and retrieval_failures ("<collection>:<reason>", where reason is
missing, embed_failed, client_unavailable or error), and context_node emits
node_finished{ok: false, retrieval_failures: [...]} so the timeline marks the node. Zero
results are not degradation — an empty rules collection is the normal steady state, and
conflating the two is what hid this for months.
Client version. qdrant-client is pinned >=1.16.0,<1.19.0: Qdrant requires a matching
major and a minor within 1 of the server, which is the qdrant/qdrant chart pinned in
terraform/modules/qdrant. Move the pin only together with the chart. Only
worker/requirements.txt reaches the cluster — ingest/ has no Dockerfile, so the
cerebrum-hub-ingest CronJob runs the worker image. The engine also caches one client
instead of building one per query, which removes three version-handshake round trips per job.
See Ingest for the embedding provider itself — there is one embedder for the
whole system, and this component does not get its own.
Impact scope. Backwards compatible and no re-embed: cerebrum_knowledge_hub is
untouched, both new payload fields default to off, and the preflight only ever adds
collections. The client pin changes nothing locally (the venv already resolved inside the
window) — it only takes effect on the next worker image build.
Tests. Unit coverage in tests/test_qdrant_provisioning.py (provisioning, idempotency,
width mismatch, the hash guard, unreachable Qdrant, a typo’d provider),
tests/test_context_engine.py (404 classification including a message-only variant, client
caching, and the “empty is not degraded” case), tests/test_context_node_health.py (the
emitted ok), and tests/test_embedding_providers.py (vector_size per provider, and that
a new collection is sized from config rather than from the vector). Run from the repo root:
knowledge-hub/packs/*/capabilities.md frontmatter-lint failures; stash and
re-run to confirm a failure is not pre-existing.
Model dispatch
Model selection is per level and driven by data, not code.config/model_catalog.yaml is the single answer to what exists, what it can do, what it costs; config/team_controls.yaml decides what this deployment may use. The loader refuses to start when either file is missing — a silently degraded catalog once disabled every provider but Anthropic while everything looked configured.
- Cost tiers are
free → cheap → mid → frontier. The user’s compute budget on/new(Free / Cost-optimized / Frontier, default cheap) resolves to a band per role (worker/worker/llm/tier_profiles.py): frontier buys a frontier architect and executor while planning and review stay cheap; Cost-optimized runs every role cheap; Free flips dispatch to the quota-paced free pool. - Capability before cost. A band applies after the level’s requirements (
require_features,min_context_tokens) and is dropped if nothing survives — a cheaper model that cannot do the work is not a saving. - Band walk on failure. The resolved band travels as ordered
ModelCandidates.call_paid_modelwalks it strongest-first — crossing providers — before any provider default; a 401/403 marks that provider dead for the rest of the call, and an unfunded account (402, or Zhipu’s 429-with-code-1113) skips the provider by name. A tool loop walks the band only until its first successful turn, after which the tool-call history is vendor-shaped and switching would corrupt it. - Effort ladder. Levels ask for
none → low → medium → high → max; adapters translate per vendor. A missing rung rounds up (exceptnone, a ceiling), and effort plus reasoning tokens are recorded per call. - Free quota ledger. Free tiers fail by refusing, so a shared Redis ledger paces requests and tokens per window across all workers (
worker/worker/llm/quota.py); a model whose window is spent is skipped, not tried. Every ledger method degrades permissive on a Redis error.
model_fallback event, is in Model routing.
Hierarchy invariants
- Sub-job nesting depth is capped at 2.
- Per-conflict escalation count maxes at 3 retries before the conflict is escalated upward.
- The L4 executor wraps token streaming with periodic heartbeats so long generations stay alive end-to-end.
- Knowledge-hub edits hot-reload into running workers via a refresher thread — no worker restart needed.