Skip to main content
worker/telemetry/snapshot.py produces the JSON the dashboard reads from /v1/command-center/snapshot. Three of its tiles were systematically overstating activity because their SQL filters were too loose; this doc records the corrected semantics so they don’t regress.

Tiles and their semantics

compute_job_flow — Stage throughput / Flows tile

  • In flight counts prompt_threads.status = 'running' only. Previously this also included queued, blocked, and awaiting_user, which inflated the count with stuck work.
  • Queued counts status in ('queued', 'blocked') as its own stage.
  • Awaiting user counts status = 'awaiting_user' as its own stage.
  • Free / Paid lane stages are now derived per-thread from the most recent job_runs.route (paid_model → paid; anything else, including not-yet-routed threads, → free). Both lanes appear in the Sankey whenever they contain work.
  • Frontend: apps/frontend/components/dashboard/widgets/job-flow.tsx carries the matching STAGE_LABEL/STAGE_COLOR entries (paid is now labeled “Paid lane” with a distinct purple tile; the previous copy-paste bug labeled it “Free lane”).
  • Workflows stat counts workflow_runs with status in ('queued', 'running', 'blocked'), completed_at IS NULL, and updated_at > now() - interval '1 hour'. The completed_at IS NULL clause prevents workflows that already reached a terminal state — but whose status was later flipped back by reconcile_workflow_status — from leaking in. The 1-hour activity window drops abandoned workflows from the number.
  • Frontend presentation (apps/frontend/components/layout/footer/activity.tsx): the category dropdown (Agents Activity / Knowledge Stats / System Specs) was replaced by gallery-style pagination dots to the right of the stat blocks. Categories auto-rotate every 15 s (infinitely); hovering the widget pauses rotation, and clicking a dot jumps to that category and resets the timer. Snapshot payload and useFooterActivity are unchanged — this is presentation-only.

compute_live_activity — Activity tracker tile (queued badge)

  • queuedTasks = prompt_threads.status in ('queued', 'blocked') AND updated_at > now() - interval '5 minutes' plus the Redis queue_depth. The recent-activity filter prevents a long-blocked prompt from showing as “queued” forever; Redis remains authoritative for the real queue.
Pricing moved out of the code (2026-09-02). It lives in config/model_catalog.yaml as pricing_per_1k per model, and the catalog refuses to load a paid chat model with no price — precisely because the old _PRICING_PER_1K dict had no OpenAI or Gemini entry and the lookup defaulted to (0.0, 0.0), so every one of those calls recorded $0.00 for months. A startup crash naming the model is better than a cost dashboard that quietly lies. worker/worker/llm/providers.py::_PRICING_PER_1K survives as the fallback for when the catalog is unreadable, and the Anthropic rows in both places carry our contracted rate (1/3.5 of the published list). cost_usd on job_runs is computed from whichever applies; the UI formats the number and never sees per-token figures.

Effort, reasoning and media (2026-09-02)

job_runs gained two columns, added through the existing idempotent alter table ... add column if not exists pattern, so no migration step: Those two answer different questions and both are needed: the strongest rung answers “how hard did this job think”, the sum answers “what did that cost”. Reporting only the last rung would hide a single expensive level; reporting only the total hides whether the spend was asked for or incidental. Folded by _fold_reasoning in worker/worker/main.py. Three more fields ride in the audit log rather than as columns, because they are per-call rather than per-job:
  • media_refs — modality, media_id, MIME, size and a content hash for every attachment. Never the bytes: text redaction cannot touch a photograph, so the trail records what was sent by reference and leaves the content in the media store where lifecycle rules expire it. A part that was NOT sent (because the model could not read it) is recorded with sent: false — the call succeeded, so this is the only trace that the image never reached a model.
  • audio_seconds — from a transcription call, taken from the provider’s own verbose_json duration. Speech is billed and rate-limited by audio time, not tokens, and a client-reported duration would be a claim rather than a measurement.
  • characters — from a speech call. Every TTS vendor bills per character.

Provider labels

_normalize_provider used to fold Anthropic into "claude" — a display name the backend never records — and mapped anything unrecognized to its bare lowercase name. Once routing opened to eight providers, the Model Mix tile started receiving "deepseek", had no colour for it, and rendered the row with undefined. It now emits the backend’s own keys (anthropic, deepseek, zhipu, moonshot, minimax, groq, openrouter, gemini, openai, ollama), the frontend covers the full set, and providerColor() falls back to grey so a provider added to the catalog appears immediately rather than invisibly.

Impact scope

All changes are read-side / cosmetic. The two new job_runs columns are added idempotently at startup, so there is no migration step and an older worker reading the table is unaffected. A worker restart picks up catalog pricing on the next job; the telemetry API picks up the new tile queries on its next snapshot. One behaviour change worth naming: rows written before this change have provider = "claude" if anything wrote that label, and the frontend no longer has a colour keyed on it — those rows render grey rather than pink. The worker itself always recorded "anthropic", so this affects only the projection.

Tests

Touched code is covered by the existing telemetry tests:
Manual verification:
  • Run select count(*) from prompt_threads where status = 'running' against the dev DB and confirm the “In flight” tile matches.
  • Trigger a paid-lane job (routing_hint=paid_model) and confirm the computed cost_usd is ~3.5× lower than a hand-multiplication of the old (0.015, 0.075) rates.