Skip to main content

What it does

  • Authenticates requests via Authorization-header API tokens against team policy (per-user and per-org rate limits; secrets/PII redaction at ingress).
  • Persists projects and workflows.
  • Enqueues every chat submission as a Redis job before any execution — the queue-first invariant.
  • Streams job events to clients over Server-Sent Events.
  • Routes non-LLM intents (rules, facts, commands), approvals, and Slack events.
  • Serves the model catalog with live free-tier quota, and the knowledge-hub query/edit surface behind the KB MCP server and the Atlas.

Endpoints

The full catalog — jobs (submit / status / stream / resume / cancel), workflows and prompts, projects, approvals, capabilities, work items, Slack, Atlas, and knowledge — is maintained in one place: HTTP API & event stream. The core ingest loop: Legacy GET|POST /v1/ollama/* probe/wake endpoints still exist for the retired self-hosted stack, which nothing routes to on purpose anymore.

Slack ingestion

POST /v1/slack/events (event subscriptions) and POST /v1/slack/interactions (approval and question buttons) make Slack a two-way front end. The load-bearing rules:
  • Signature verification is the only inbound auth — an unset SLACK_SIGNING_SECRET rejects every request, silently.
  • Triage decides act vs. ignore. Amos acts on a DM, an @mention, or a reply inside a thread it already holds a conversation in (slack_threads is the lookup; a top-level channel message is never a follow-up). Everything else is dropped without a word: its own posts (bot_id present), Slackbot posts (which carry no bot_id — they post as a user), messages with attachments (a canvas/screenshot is boilerplate around a link; the content reaches Amos through the ingest connector instead), any subtype (message_changed would re-run a whole turn on a typo fix), duplicate deliveries, and authors outside the workspace. Drop reasons are logged at debug and name the exact rule.
  • One turn at a time per thread. A second message while the first runs gets “still working on your previous message”. Thread history is the worker’s usual recency window, not the full Slack transcript.
  • Thread identity is a table, not metadata. slack_threads (primary key channel, thread_ts → workflow_id, user_id, last_job_id) exists because claiming a thread must be atomic — two near-simultaneous messages must not split the history — and because workflow metadata is rewritten on every append. The reverse user lookup uses a deliberately non-unique partial index: the column is nullable and duplicates are legal, so a unique index would fail ensure_schema() and prevent the worker from booting.
  • Derived, not configured. The workspace id and the bot’s own user id come from a cached auth.test call — configuring them separately only creates drift that fails closed (every message refused as “outside the allowed workspace”). SLACK_TEAM_ID remains an override. The signing secret is the one credential nothing can derive; everything inbound 401s until it is set. Credentials live in one Secrets Manager secret (cerebrum/connectors/slack); the setup runbook is docs/slack-ingestion.md.

KB retrieval embeds elsewhere

(2026-09-03) POST /v1/kb/query — the endpoint behind the KB MCP tool, the capability acceptance check and the frontend’s sources panel — no longer embeds its own query vector. It asks telemetry-api’s POST /v1/kb/embed and then searches Qdrant as before (routes::embed_via_telemetry). It used to embed against Ollama on the GPU node, and would wait up to ten minutes for a wake. That is why the task modal showed “Waking GPU… ~2:36 remaining” while someone read an answer about bloom filters: opening a job panel started a g5.xlarge to render a sidebar. mother-ai cannot embed against a hosted provider itself, for two reasons that are both about not spreading credentials:
  • The key is in Secrets Manager and reading it needs an AWS SDK this service has no other use for. Injecting it as an env var was ruled out by a principle already written into terraform/modules/provider-secrets: a provider key in state is a provider key in every plan output and every state backup.
  • OpenAI geo-blocks ap-east-1, so the call needs the egress proxy that telemetry-api already configures.
The deeper reason is that there must be exactly one embedder: the indexer and every querier have to use the same model at the same width forever, and a second implementation in another language is a second place for that to drift into a different vector space — a failure that returns confident nonsense rather than an error. EMBEDDING_PROVIDER=ollama still takes the local path, so a rollback is one env var (plus a re-embed, since the vectors move with it).

Identifiers: wf-N and pt-N come from Postgres

(2026-09-03) Workflow and prompt ids are allocated by mother-ai/src/ids.rs from two Postgres sequences, cerebrum_workflow_id_seq and cerebrum_prompt_id_seq. They used to come from a Redis INCR on a key namespaced by instance_id — which config::derive_instance_hash derives from the directory holding the control-plane config. That made the numbering a function of the image layout, and it broke exactly that way: adding COPY config /app/config to the Dockerfile turned a failing canonicalize() into a succeeding one, the hash moved from sha256("config") to sha256("/app/config"), the counter restarted at 1, and a new workflow was handed wf-9 — an id belonging to a workflow from five months earlier. Since prompt_threads.workflow_id is the only link between a workflow and its prompts, the new job and a stranger’s April prompt rendered as one two-step workflow: the graph showed a level that never ran, and the detail panel showed an answer to a question nobody had just asked. The id namespace now lives where the rows it names live. A counter in the cache can only ever be a guess about what the durable store already contains — it restarts on a flush, an eviction, or a key rename, and nothing notices until two eras of data are stitched together in one view. Two properties worth knowing:
  • Self-repairing. On the first allocation in a process the sequence is aligned past the highest id any of prompt_threads, workflow_runs and job_runs has already issued, and alignment only ever moves forward. A database restored from a dump, or a counter that already drifted, converges without a migration.
  • No fallback when Postgres is configured. A sequence failure returns an error rather than reaching for the Redis counter, because quietly reaching for that counter is the failure being removed. Redis remains the allocator only when no POSTGRES_DSN is set at all — a local checkout with nothing to collide with.
instance_id still namespaces the job queue key and the rate-limit counters. Those are pinned explicitly by Terraform (JOB_QUEUE_KEY on all three services), so a path change re-namespaces nothing that outlives a process.

SSE invariant

The stream handler must replay the durable backlog before joining the pubsub stream so that a reconnecting client never misses events.

Errors

Errors include job and request correlation metadata so clients can trace any failure back to the job and the original request.