Shared base
Every connector subclassesBaseConnector (ingest/connectors/base.py) and implements one method:
- Deterministic identity —
kh-idiskh-<26-char Crockford32>derived from a stablesource_id(SHA-256, first 16 bytes encoded). Same upstream record always yields the same id, so updates rewrite the same chunk rather than spawning duplicates. - Idempotent writes — when an existing file’s canonical frontmatter + body match the new render, the write is skipped entirely. Existing
idandcreated_atare preserved across rewrites. - Cursor persistence — the Postgres
connector_statetable (one row per connector) stores a JSON cursor so the next run picks up where this one left off. The cursor is rolled back if the reindex fails:run_connector.main()snapshots it before the run and restores it whenreindex_pathsraises. Without that, a failed reindex left the cursor reading “processed through here” while the records were never in Qdrant — and because the hub mount is ephemeral, the rendered files were gone too. Those records would never be fetched again: silently absent from retrieval, with no error after the failing run. - Schema-compliant frontmatter — every emitted document validates against
knowledge-hub/.kb/schema/frontmatter.schema.json(required fields, sensitivity enum,source_hashformat).
Trigger model
All three are time-driven CronJobs, no webhooks. The cursor does deduplication so polling often when there’s nothing new is cheap.
Each pod’s lifecycle:
- Init container clones the curated hub into an emptyDir (same pattern as
amos-hub-ingest). - Main container runs
python -m ingest.run_connector --connector <name>. - The connector writes new / changed markdown into the emptyDir under
sources/<service>/.... - The same entrypoint immediately calls
ingest_hub.reindex_paths()on the changed files, so Qdrant has the new chunks before the pod exits. - Pod exits, emptyDir wiped. Durable state lives in Postgres (cursor) and Qdrant (content).
Notion connector
What it pullsPOST /v1/searchwithfilter={object:page}sorted descending bylast_edited_time.- For each page newer than the cursor:
GET /v1/blocks/{id}/childrenrecursively to build the block tree.
Rich-text annotations stack
code → bold → italic → strikethrough → underline (underline renders as italic; markdown has no native underline). Links wrap the already-annotated text.
Cursor: {"last_edited_time": "2026-05-12T10:00:00Z"}. Search is sorted desc, so the iterator breaks on the first page older than the cursor.
Output: sources/notion/<workspace_slug>/<page-slug>.md. Untitled pages fall back to the first 12 chars of the page id as a slug.
source_id: notion:page:<page-id> — stable across edits.
Use cases
- “What did we decide about retention in the Atlas spec?” — product docs in Notion.
- Vendor and project runbooks that don’t get migrated to the curated hub.
- OKRs, weekly reviews, meeting notes.
Slack connector
What it pulls (per channel in the allowlist)conversations.historywitholdest=<cursor>, paginated vianext_cursor. Returns messages descending byts; the connector sorts ascending before yielding.- For each message where
ts == thread_tsANDreply_count > 0:conversations.repliesto fetch the thread tail. Replies already in history are deduplicated byts.
conversations.history returns not_in_channel unless the bot is a member, even
for public channels. The token carries channels:history and channels:read but
not channels:join, so the app cannot add itself — someone has to run
/invite @<bot> in each channel, or channels:join has to be added to the Slack
app and the app reinstalled.
A channel the bot cannot read is skipped, not fatal. SlackApiError carries
Slack’s machine-readable error, and not_in_channel / channel_not_found /
is_archived are logged with the remedy and skipped, leaving that channel’s
cursor untouched so nothing is marked read that wasn’t. Any other Slack error
still fails the run.
A run that reads zero channels fails. Degrading per-channel is correct;
degrading to nothing is not. This is not hypothetical: the bot was removed from
all three configured channels, every run aborted on the first one, and because
the only symptom was a failing CronJob nobody was watching, Slack ingestion was
dead for weeks while the index kept serving a static export.
File attachments — canvases carry the meeting notes
Slack huddle notes are posted by Slackbot as canvases (filetype: quip,
application/vnd.slack-docs) with the message text empty. Rendering only the
filename meant the most substantive content in a channel — actual meeting
discussion — was invisible to retrieval: #all-gdi-labs had 44 such canvases
(~208 KB, May–Sep 2026) indexed as nothing but a slugified link.
Text-bearing attachments now have their bodies fetched and embedded. Governed by
TEXT_FILETYPES / TEXT_MIMETYPES (canvas, post, text, markdown, text/*);
images, PDFs, video and archives stay links, because there is no extractor for
them and a binary blob in a day-file degrades retrieval rather than helping it.
Capped at MAX_FILE_BYTES (256 KB) per file so one attachment can’t dominate a
day-file.
- Needs the
files:readscope. Without it every fetch returns 403 and each file degrades to today’s link-only rendering — the feature is inert, not broken.SLACK_INGEST_FILE_CONTENTS=0disables it outright. - Sensitivity is inherited from the channel, exactly like the surrounding messages. A file is not treated as more public than the conversation it was posted in.
- The label prefers the file’s
titleover itsname: huddle notes arrive as_headphones__Huddle_notes__8_31_26_in___C0B33NXGBSA_, so the slug was what got indexed and a search for “huddle notes August” matched nothing. - Canvas serialization is not a documented Slack format, so the converter is deliberately conservative: strip tags when the body looks like HTML, otherwise pass through. The fetch logs the real content type so this can be sharpened from a live run.
ts is converted to UTC and floored to YYYY-MM-DD. Messages with the same date land in one _DayBucket; the connector yields one record per bucket. last_ts per bucket — the highest ts in it — advances cursor[channel_id].
Rewind window — rewind_days (default 1)
Day-files for the past N days are treated as still mutable. Each run computes effective_oldest = min(cursor, start_of_today_utc - rewind_days) and refetches everything in that window, then re-renders the affected day-files. This guarantees that a tick mid-day (or right after midnight) doesn’t truncate today’s file to “only the messages since the last tick” — without the window, every same-day rerun would have replaced today’s day-file with only the new messages.
The cursor still advances to the latest ts seen, so days outside the window stay immutable and aren’t refetched. Bumping rewind_days widens the late-edit catch (Slack lets users edit messages weeks later) at the cost of one extra full-day fetch per channel per run.
Thread stitching
After all messages for a day are collected, _stitch_threads groups them: root messages (ts == thread_ts) become parents; replies (ts != thread_ts) attach under their parent. Each thread renders as one ## HH:MM — <topic-snippet> heading in the day file.
Text transforms (in order)
- User mentions
<@U001>→@<display-name>. Resolved lazily viausers.info, cached for the run. Unknown users render as@U001. - Channel mentions
<#C111|design>→#design. Falls back to#C111if Slack didn’t include the name. - Links
<https://x.test|label>→[label](https://x.test); bare<https://x.test>→https://x.test. html.unescapecollapses&,<, etc.
/archives/<channel>/p<ts-no-dot>) so any entry links back to the live thread. Files cover both permalink and url_private. Reactions are an italicized inline footer.
Per-channel policy (SlackChannelConfig)
Each channel carries its own scope, sensitivity (public | partner | internal | secret), and tags. So #general can be internal while #nuts-bleu-qa is partner and #pike-tapio-launch-partners is also partner.
Cursor: {channel_id: latest_ts_string}. Each channel advances independently.
Output: sources/slack/<channel_name>/<YYYY-MM-DD>.md.
source_id: slack:<channel_id>:<YYYY-MM-DD> — same channel + same day always yields the same kh-id, so re-runs are idempotent.
Use cases
- “Why did we pick UUID v7 for the workflows table?” — the decision conversation in
#general. - Async standup context: who’s blocked on what, surfaced when an agent picks up a half-finished task.
- For agents: cross-referencing diary entries with the Slack thread that triggered them — same
task_idtag.
ingest/connectors/slack_backfill.py, handles re-seeding a channel from a workspace export when bot access only landed recently. The daily-driver path is the connector above.
Not the only Slack path any more. This connector is a scheduled pull into the hub, and it stays exactly that. Since 2026-09-05 Slack is also a live conversational surface — a DM or @Amos becomes a turn on a workflow, and the answer comes back in the thread. That is a separate mechanism with its own credential, endpoint and storage; see Slack Conversation. Its mrkdwn converter is the deliberate inverse of this connector’s transform_slack_text, and the two are worth keeping in lockstep.
GitHub connector
What it pulls- For each repo in the allowlist:
git clone --depth 1 --branch <branch> --single-branchinto a temp dir (auth via token-in-URL when configured). git rev-parse HEAD→ current commit SHA.- If
cursor[<owner/repo>] == sha: skip the whole repo. Repos that haven’t moved cost one clone + onerev-parse.
os.walk over the cloned tree (not pathlib.glob, which hides dotfiles on Python 3.12 so .git/** exclusions would silently miss). For each file:
- Compute repo-relative posix path.
- Match against
exclude_globs→ skip if any match. - Match against
include_globs→ skip if none match. stat().st_size > max_file_bytes(default 256 KB) → skip.read_text(encoding="utf-8")—UnicodeDecodeErroris treated as binary and skipped.- Emit a record.
** means what users expect: **/node_modules/** correctly excludes apps/frontend/node_modules/anything. (fnmatch treats ** as *, which is wrong for path globs.)
File rendering
{<owner/repo>: <commit-sha>}.
Output: sources/github/<owner>/<repo>/<relative-path>.md.
source_id: github:<owner>/<repo>:<path> — deliberately omits the SHA. A new commit that edits lib.py keeps the same source_id, so the same kh-id, so the chunk updates rather than spawning a duplicate. The SHA is in source_url for traceability but not in identity.
Use cases
- “Find the function that resolves Mother AI URLs.” — literal-jargon queries where dense embeddings blur token boundaries.
- “Where do we set the Qdrant collection name?” — config-string lookup across repos.
- Cross-repo references: ingest
amosandamos-templatestogether so workers can reason about templates while looking at how they get applied. - A motivating workload for the hybrid retrieval (BM25 + dense + RRF) work tracked next: identifier-style code tokens match literally, which is exactly where pure dense embeddings underperform.
Configuration
Every CronJob defaults off. Per-connector knobs live interraform/workloads/environments/dev/terraform.tfvars:
*_auth_secret_id points at an AWS Secrets Manager secret resolved at runtime via IRSA. Per-channel and per-repo policy ships as JSON in the env var (parsed by ingest/run_connector.py).
Manual trigger
Invariants
- Source-system credentials never leave AWS Secrets Manager. The Python entrypoint resolves them at startup; no token lives in a manifest or env-fallback in production.
- Cursor + Qdrant carry the durable state. A connector pod’s emptyDir is wiped on exit; this is intentional. The cursor in Postgres prevents re-fetching upstream; Qdrant payload carries chunk content.
- Sensitivity is per-record, not per-connector. A Slack channel can be
partner; a GitHub repo can beinternal; a Notion page tier is the connector default. Retrieval-time tier enforcement (when it lands) filters on this. - Idempotent writes. The same upstream record on a re-run produces the same kh-id and skips the write when content matches. Reseeding a connector is safe.