Skip to main content

What it enables

  • Decouple agent flow shape from worker code. New scenarios — code, writing, research, security, Q&A — ship as YAML + prompt builders + (optional) rubric markdown. No worker deploy required for a new flow once the relevant runners exist.
  • Per-job scenario dispatch. Each scenario YAML declares the intent_route values it serves (code_build → project, document_writing → document/writing, quick_qa → task/qa, …). The worker compiles every loaded scenario once at startup and picks one per job from (in priority): payload.metadata["scenario_id"] → payload.metadata["intent_route"] mapped via the loader → AMOS_DEFAULT_SCENARIO. The command widget’s task type sets intent_route; the future universal-input classifier will set it (or pin scenario_id directly) by classifying the prompt.
  • Per-level model pinning. Each level resolves its own ModelConstraint through worker/worker/llm/resolver.py, so an L4 executor can be pinned to Claude Opus while a Q&A scenario runs the whole thing on the free tier.
  • Pluggable verifiers (subprocess + LLM judge + schema + MCP). Verifiers can run on prompt-kind levels (e.g., an llm_judge rubric on l3_accept) or on a dedicated verifier-kind level that gates the next branch.
  • First-class DAG branching. Edges carry a typed boolean DSL (verdict == 'PASS' && !escalate_now); the compiler evaluates them at routing time. fan_out parents run children in parallel; aggregators (majority_pass, weighted_score, all_pass) merge results.
  • Agentic file-tool loop. kind: tool_loop levels (l4_execute, l2_escalate) mutate the project workspace through tool calls the worker executes — surgical edit_files on the scaffolded template, not a full-file Markdown blob. The workspace filesystem is the shared state; levels exchange paths + summaries.
  • Human approval gates. A kind: human_approval level pauses the graph via LangGraph’s interrupt(), packages a question for the approver, and resumes with their reply as the gate-readable verdict.
  • Budget kill switch + user Stop. Scenario-wide caps (cost / tokens / duration) and a user-initiated Stop both set aborted_reason, which short-circuits subsequent levels with synthetic skipped states and routes the gate straight to END. The run ends scenario_finished status="aborted" (a cap) or "cancelled" (a Stop); a cancelled run discards partial work — no commit/push. Stop is a Redis flag (amos:job:{id}:cancel) the worker polls between levels and per tool-turn via EventBus.is_cancelled().
  • Per-tenant overrides. A partial YAML stored in Postgres deep-merges over the base scenario at orchestrator construction time, so an org can tweak budgets / constraints / verifier weights without forking the base file.

Module layout

All paths under worker/worker/graph/scenarios/:
  • schemas.py — Pydantic models with strict (extra="forbid") validation. Scenario, Level, Edge, Gate, Budget, ModelConstraint, VerifierRef, FanInRef. Composite primary-key validation (unique level ids, edges target existing levels or the __end__ sentinel, mixed-tier allows rejected).
  • state.py — ScenarioState TypedDict threaded through the LangGraph. Explicit top-level fields for every value the gate DSL or runners read: verdict, escalate_now, confidence, aborted_reason, budget_exceeded, per-level entries under levels: dict[str, LevelState].
  • scratchpad.py — Redis-backed scenario-scoped KV (amos:job:{job_id}:scratchpad). Available to every level via the scratchpad.{get,put,list} intra-tools.
  • expr.py — Recursive-descent parser for the edge DSL. Supports paths (verifier.build.passed), string/number/bool literals, ==/!=/<,<=,>,>=, !, &&, ||, parens. Missing paths resolve to a fail-closed UNDEFINED sentinel. Loader compiles every Edge.when at load time.
  • loader.py — ScenarioLoader reads *.yaml from knowledge-hub/scenarios/. refresh() swaps the registry atomically and preserves the old version on any validation error. with_overrides({tenant_id: partial_yaml}) deep-merges tenant tweaks. Scenario.checksum is the sha256 of the canonical YAML, recorded on each ScenarioResult for reproducibility.
  • compiler.py — ScenarioCompiler turns a Scenario into a LangGraph StateGraph. One node per level; non-conditional edges via add_edge; multi-branch edges via add_conditional_edges with a router that evaluates the edge DSL and emits gate_decision events.
  • executor.py — ScenarioExecutor dispatches each level to its kind-specific runner. Wraps every invocation with level_started / level_finished events alongside the legacy node_started / node_finished pair (back-compat through one release).
  • runners.py — PromptLevelRunner, ToolLoopLevelRunner, VerifierLevelRunner, FanOutLevelRunner, HumanApprovalLevelRunner. See Scenario execution for the per-runner semantics. ToolLoopLevelRunner drives the kind: tool_loop levels (l4_execute, l2_escalate): an agentic file-tool loop where the model mutates the workspace through read_file / edit_file / write_file / list_dir / delete_file / finish tools — the filesystem is the artifact, downstream levels receive paths, not file bodies.
  • prompts.py — Registered prompt builders for every scenario level (l1_team_lead_brief, writing_outline, security_triage, qa_answer, etc.). Builders return a BuiltPrompt { body, response_kind, handoff_* } and read prior-level outputs from state["levels"]. code_build’s L1 emits a Markdown brief and L3 emits Markdown prose + a trailing ```control block carrying the verdict (worker/runtime/control_block.py) — not a JSON AgentMessage envelope; the runner runs a bounded repair-retry when the verdict won’t parse.
  • worker/runtime/file_tools.py — the six sandboxed file tools + the native (Anthropic) and textual transports the tool loop dispatches.
  • aggregators.py — MajorityPassAggregator, WeightedScoreAggregator, AllPassAggregator. Aggregator name is referenced from the parent Level.fan_in.aggregator field.
  • tenant_overrides.py — TenantOverrideProvider Protocol with two impls: InMemoryTenantOverrideProvider (tests) and PostgresTenantOverrideProvider (production). DB schema is on worker/worker/storage/postgres_repo.py under tenant_scenario_overrides.
  • replay.py — CassetteProviders + ReplayHarness. Drives a scenario through recorded model responses (no live calls) and asserts no verifier or rubric-score regression beyond the scenario’s regression_thresholds. Runs in CI via .github/workflows/scenario-replay.yml.
  • orchestrator.py — ScenarioOrchestrator is the public entry the worker calls. Same run(payload, routing_decision, event_bus) -> ScenarioResult contract the legacy HierarchyOrchestrator exposed.
The verifier framework lives at worker/worker/runtime/verifiers/ and is consumed by both PromptLevelRunner (when a prompt level declares verifiers) and VerifierLevelRunner. Built-ins: build (subprocess install + build), typecheck, lint, unit_test, schema (JSON Schema + Pydantic), llm_judge (rubric-scored), semgrep, gitleaks.

Lifecycle events

Worker emits both the legacy and the scenario lifecycle events. Frontend renders the new events via apps/frontend/components/workflow/scenario-run-timeline.tsx. Payload contracts in worker/worker/runtime/agent_messages.py. TypeScript mirror in apps/frontend/lib/types/agent-event.ts. The full catalog with semantics lives in HTTP API & event stream.

Runner semantics

ScenarioExecutor dispatches each level to its kind runner; the per-kind behavior every consumer should know:
  • prompt (PromptLevelRunner) — resolves the prompt builder, resolves the model (emitting the per-level routing_decision), calls under a heartbeat for long generations, records cost/tokens, then post-processes: l1_brief emits Markdown directly; l3_accept/fact_check extract a PASS/FAIL verdict from a trailing ```control block (with one bounded repair-retry when it won’t parse); terminal-output levels set the final output. A builder may return a static_output — a level composed entirely in code (the code_build L1 brief built from the user’s confirmed features is the shipped case) — which sets deterministic: true and makes no model call.
  • tool_loop (ToolLoopLevelRunner) — the agentic file loop over read_file / write_file / edit_file / list_dir / delete_file / finish, up to max_iterations (default 40). Paid transport returns native tool-use blocks; the textual ```tool protocol over the free tier is available only when the first paid turn fails and the level allows it. The run’s Stop flag is checked per turn. The band walk applies only until the first successful turn — the tool-call history is vendor-shaped after that.
  • verifier (VerifierLevelRunner) — no file-apply step (the loop already wrote to disk); builds a VerifierContext against the workspace and runs the declared verifiers (serial / parallel / race). Every failed verifier’s reasons are stashed as build_feedback for the next builder attempt. code_build runs build plus the deterministic route_reachability gate (every App Router route navigable from /; orphan and misplaced routes flagged).
  • human_approval / human_feedback (HumanApprovalLevelRunner) — parks the graph via LangGraph interrupt(), packages the question (approval or the universal feedback contract), and resumes with the reply folded back per the level’s metadata.assign (intent_route chaining or a state key).
  • fan_out/aggregator (FanOutLevelRunner) — runs children in parallel through the same executor, merges their level states, and applies the named aggregator; the parent’s post-merge budget check is what stops a fan-out from slipping past the cap.

State persistence rules

LangGraph persists a state channel only when a node returns it in its delta — a channel mutated in place but dropped from the return value silently reverts on every re-hydration (each ask_user resume, worker restart, and re-enqueue). This produced two real failures: a retry counter that reset mid-run (five build-verify laps against a budget of three), and an audit list that grew checkpoint writes quadratically (one run wrote 338 MB of checkpoints). Both fixes generalize into rules for new state:
  • If a node changes a mutable channel, return it — trackers return (escalate, tracker) so a caller cannot forget.
  • Full audit entries (whole prompt, raw response, tool calls) go to the model_call_audit Postgres sink, one row per call; graph state keeps only metadata, and the orchestrator reassembles the full list post-run.
  • A handle rather than data (an event bus, a sink, a client) stays off ScenarioState entirely and is reached through a ContextVar — declaring it on state would make it a checkpointed channel.

Bundled scenarios

Rubrics live alongside scenarios in knowledge-hub/rubrics/: architecture_review_v1, fact_check_v1, rigor_v1.

Authoring a scenario

A scenario ships as YAML + prompt builders + (optional) rubric markdown. Rules that aren’t obvious from the schema:
  • Declare examples: (3–5 representative prompts) if the scenario should be reachable from the universal input. The intent classifier matches submissions against this catalog; admin-only scenarios (e.g. code_build_strict) deliberately omit examples to stay out of it. Add the examples in the same PR that adds the scenario.
  • intent_routes: are first-come. Two scenarios claiming the same route is a hard load error. The task route is reserved for the tier-3 classifier rule — a plain “task” with no explicit route goes to intent_classifier, which picks by examples; direct callers should pin scenario_id or use a specific route.
  • Chain depth is capped at 3, and the classifier counts as one hop.
  • Non-conversational work types (rule, fact, planner, metric) bypass the orchestrator entirely.

CI gate

.github/workflows/scenario-replay.yml runs on every PR touching knowledge-hub/scenarios/**, knowledge-hub/rubrics/**, or worker/worker/graph/scenarios/**. It loads each scenario’s golden fixtures from knowledge-hub/scenarios/<id>/golden/*.json and replays them through ReplayHarness against CassetteProviders. The job fails on any verifier or rubric regression beyond declared thresholds.

Impact scope

  • Worker: hot path for every paid job goes through ScenarioOrchestrator. Legacy HierarchyOrchestrator is deleted; the wrapping WorkerOrchestrator is unchanged.
  • Frontend: JobDetailPanel mounts ScenarioRunTimeline for every job. Renders nothing on legacy backlog events for backwards compat.
  • Postgres: one new table tenant_scenario_overrides (composite PK on tenant_id, scenario_id).
  • Defaults: code_build is the default scenario; override via the AMOS_DEFAULT_SCENARIO env var or per-job scenario_id metadata.
  • Backwards compat: legacy node_started / node_finished events still fire alongside the new level_* events for one release.

Tests

All tests live at repo root tests/. The scenario runtime suite:
  • test_scenarios.py (28) — schemas, loader, scratchpad.
  • test_scenario_expr.py (25) — gate DSL.
  • test_verifiers.py + test_verifiers_phase2.py (55) — verifier framework + built-ins.
  • test_scenario_compiler.py (10) — compiler topology + end-to-end with mock runners.
  • test_scenario_prompts.py (10) — prompt-builder parity.
  • test_scenario_orchestrator.py (9) — production orchestrator with fake providers.
  • test_scenario_events.py (13) — Phase-4 lifecycle event payloads.
  • test_scenario_phase5.py (12) — fan_out + aggregators + budget kill switch.
  • test_scenario_phase6.py (16) — checksum, hot-reload, tenant overrides, replay harness.
  • test_tenant_overrides.py (12) — DB-backed tenant overrides.
  • test_model_resolver.py (15) — per-level model pinning.
  • test_document_writing.py (15), test_quick_qa_and_research_brief.py (12), test_security_review.py (15) — bundled scenarios end-to-end.
Run from repo root with the worker venv:
Suite status as of 2026-05-13: 502 passing, 9 unrelated pre-existing failures (test_provider_registry.py, test_scheduler_producer.py, test_context_engine.py — all WorkerConfig.__init__ drift not introduced by the scenario work).