What it enables
- Decouple agent flow shape from worker code. New scenarios — code, writing, research, security, Q&A — ship as YAML + prompt builders + (optional) rubric markdown. No worker deploy required for a new flow once the relevant runners exist.
- Per-job scenario dispatch. Each scenario YAML declares the
intent_routevalues it serves (code_build→project,document_writing→document/writing,quick_qa→task/qa, …). The worker compiles every loaded scenario once at startup and picks one per job from (in priority):payload.metadata["scenario_id"]→payload.metadata["intent_route"]mapped via the loader →AMOS_DEFAULT_SCENARIO. The command widget’s task type setsintent_route; the future universal-input classifier will set it (or pinscenario_iddirectly) by classifying the prompt. - Per-level model pinning. Each level resolves its own
ModelConstraintthroughworker/worker/llm/resolver.py, so an L4 executor can be pinned to Claude Opus while a Q&A scenario runs the whole thing on the free tier. - Pluggable verifiers (subprocess + LLM judge + schema + MCP). Verifiers can
run on prompt-kind levels (e.g., an
llm_judgerubric onl3_accept) or on a dedicated verifier-kind level that gates the next branch. - First-class DAG branching. Edges carry a typed boolean DSL
(
verdict == 'PASS' && !escalate_now); the compiler evaluates them at routing time.fan_outparents run children in parallel; aggregators (majority_pass,weighted_score,all_pass) merge results. - Agentic file-tool loop.
kind: tool_looplevels (l4_execute,l2_escalate) mutate the project workspace through tool calls the worker executes — surgicaledit_files on the scaffolded template, not a full-file Markdown blob. The workspace filesystem is the shared state; levels exchange paths + summaries. - Human approval gates. A
kind: human_approvallevel pauses the graph via LangGraph’sinterrupt(), packages a question for the approver, and resumes with their reply as the gate-readableverdict. - Budget kill switch + user Stop. Scenario-wide caps (cost / tokens /
duration) and a user-initiated Stop both set
aborted_reason, which short-circuits subsequent levels with synthetic skipped states and routes the gate straight toEND. The run endsscenario_finished status="aborted"(a cap) or"cancelled"(a Stop); a cancelled run discards partial work — no commit/push. Stop is a Redis flag (amos:job:{id}:cancel) the worker polls between levels and per tool-turn viaEventBus.is_cancelled(). - Per-tenant overrides. A partial YAML stored in Postgres deep-merges over the base scenario at orchestrator construction time, so an org can tweak budgets / constraints / verifier weights without forking the base file.
Module layout
All paths underworker/worker/graph/scenarios/:
schemas.py— Pydantic models with strict (extra="forbid") validation.Scenario,Level,Edge,Gate,Budget,ModelConstraint,VerifierRef,FanInRef. Composite primary-key validation (unique level ids, edges target existing levels or the__end__sentinel, mixed-tier allows rejected).state.py—ScenarioStateTypedDict threaded through the LangGraph. Explicit top-level fields for every value the gate DSL or runners read:verdict,escalate_now,confidence,aborted_reason,budget_exceeded, per-level entries underlevels: dict[str, LevelState].scratchpad.py— Redis-backed scenario-scoped KV (amos:job:{job_id}:scratchpad). Available to every level via thescratchpad.{get,put,list}intra-tools.expr.py— Recursive-descent parser for the edge DSL. Supports paths (verifier.build.passed), string/number/bool literals,==/!=/<,<=,>,>=,!,&&,||, parens. Missing paths resolve to a fail-closedUNDEFINEDsentinel. Loader compiles everyEdge.whenat load time.loader.py—ScenarioLoaderreads*.yamlfromknowledge-hub/scenarios/.refresh()swaps the registry atomically and preserves the old version on any validation error.with_overrides({tenant_id: partial_yaml})deep-merges tenant tweaks.Scenario.checksumis the sha256 of the canonical YAML, recorded on eachScenarioResultfor reproducibility.compiler.py—ScenarioCompilerturns aScenariointo a LangGraphStateGraph. One node per level; non-conditional edges viaadd_edge; multi-branch edges viaadd_conditional_edgeswith a router that evaluates the edge DSL and emitsgate_decisionevents.executor.py—ScenarioExecutordispatches each level to its kind-specific runner. Wraps every invocation withlevel_started/level_finishedevents alongside the legacynode_started/node_finishedpair (back-compat through one release).runners.py—PromptLevelRunner,ToolLoopLevelRunner,VerifierLevelRunner,FanOutLevelRunner,HumanApprovalLevelRunner. See Scenario execution for the per-runner semantics.ToolLoopLevelRunnerdrives thekind: tool_looplevels (l4_execute,l2_escalate): an agentic file-tool loop where the model mutates the workspace throughread_file/edit_file/write_file/list_dir/delete_file/finishtools — the filesystem is the artifact, downstream levels receive paths, not file bodies.prompts.py— Registered prompt builders for every scenario level (l1_team_lead_brief,writing_outline,security_triage,qa_answer, etc.). Builders return aBuiltPrompt { body, response_kind, handoff_* }and read prior-level outputs fromstate["levels"]. code_build’s L1 emits a Markdown brief and L3 emits Markdown prose + a trailing```controlblock carrying the verdict (worker/runtime/control_block.py) — not a JSONAgentMessageenvelope; the runner runs a bounded repair-retry when the verdict won’t parse.worker/runtime/file_tools.py— the six sandboxed file tools + the native (Anthropic) and textual transports the tool loop dispatches.aggregators.py—MajorityPassAggregator,WeightedScoreAggregator,AllPassAggregator. Aggregator name is referenced from the parentLevel.fan_in.aggregatorfield.tenant_overrides.py—TenantOverrideProviderProtocol with two impls:InMemoryTenantOverrideProvider(tests) andPostgresTenantOverrideProvider(production). DB schema is onworker/worker/storage/postgres_repo.pyundertenant_scenario_overrides.replay.py—CassetteProviders+ReplayHarness. Drives a scenario through recorded model responses (no live calls) and asserts no verifier or rubric-score regression beyond the scenario’sregression_thresholds. Runs in CI via.github/workflows/scenario-replay.yml.orchestrator.py—ScenarioOrchestratoris the public entry the worker calls. Samerun(payload, routing_decision, event_bus) -> ScenarioResultcontract the legacyHierarchyOrchestratorexposed.
worker/worker/runtime/verifiers/
and is consumed by both PromptLevelRunner (when a prompt level declares
verifiers) and VerifierLevelRunner. Built-ins: build (subprocess install +
build), typecheck, lint, unit_test, schema (JSON Schema + Pydantic),
llm_judge (rubric-scored), semgrep, gitleaks.
Lifecycle events
Worker emits both the legacy and the scenario lifecycle events. Frontend renders the new events viaapps/frontend/components/workflow/scenario-run-timeline.tsx.
Payload contracts in
worker/worker/runtime/agent_messages.py.
TypeScript mirror in
apps/frontend/lib/types/agent-event.ts.
The full catalog with semantics lives in HTTP API & event stream.
Runner semantics
ScenarioExecutor dispatches each level to its kind runner; the per-kind behavior every consumer should know:
prompt(PromptLevelRunner) — resolves the prompt builder, resolves the model (emitting the per-levelrouting_decision), calls under a heartbeat for long generations, records cost/tokens, then post-processes:l1_briefemits Markdown directly;l3_accept/fact_checkextract a PASS/FAIL verdict from a trailing```controlblock (with one bounded repair-retry when it won’t parse); terminal-output levels set the final output. A builder may return astatic_output— a level composed entirely in code (thecode_buildL1 brief built from the user’s confirmed features is the shipped case) — which setsdeterministic: trueand makes no model call.tool_loop(ToolLoopLevelRunner) — the agentic file loop overread_file/write_file/edit_file/list_dir/delete_file/finish, up tomax_iterations(default 40). Paid transport returns native tool-use blocks; the textual```toolprotocol over the free tier is available only when the first paid turn fails and the level allows it. The run’s Stop flag is checked per turn. The band walk applies only until the first successful turn — the tool-call history is vendor-shaped after that.verifier(VerifierLevelRunner) — no file-apply step (the loop already wrote to disk); builds aVerifierContextagainst the workspace and runs the declared verifiers (serial / parallel / race). Every failed verifier’s reasons are stashed asbuild_feedbackfor the next builder attempt.code_buildrunsbuildplus the deterministicroute_reachabilitygate (every App Router route navigable from/; orphan and misplaced routes flagged).human_approval/human_feedback(HumanApprovalLevelRunner) — parks the graph via LangGraphinterrupt(), packages the question (approval or the universal feedback contract), and resumes with the reply folded back per the level’smetadata.assign(intent_routechaining or a state key).fan_out/aggregator (FanOutLevelRunner) — runs children in parallel through the same executor, merges their level states, and applies the named aggregator; the parent’s post-merge budget check is what stops a fan-out from slipping past the cap.
State persistence rules
LangGraph persists a state channel only when a node returns it in its delta — a channel mutated in place but dropped from the return value silently reverts on every re-hydration (eachask_user resume, worker restart, and re-enqueue). This produced two real failures: a retry counter that reset mid-run (five build-verify laps against a budget of three), and an audit list that grew checkpoint writes quadratically (one run wrote 338 MB of checkpoints). Both fixes generalize into rules for new state:
- If a node changes a mutable channel, return it — trackers return
(escalate, tracker)so a caller cannot forget. - Full audit entries (whole prompt, raw response, tool calls) go to the
model_call_auditPostgres sink, one row per call; graph state keeps only metadata, and the orchestrator reassembles the full list post-run. - A handle rather than data (an event bus, a sink, a client) stays off
ScenarioStateentirely and is reached through a ContextVar — declaring it on state would make it a checkpointed channel.
Bundled scenarios
Rubrics live alongside scenarios in
knowledge-hub/rubrics/:
architecture_review_v1, fact_check_v1, rigor_v1.
Authoring a scenario
A scenario ships as YAML + prompt builders + (optional) rubric markdown. Rules that aren’t obvious from the schema:- Declare
examples:(3–5 representative prompts) if the scenario should be reachable from the universal input. The intent classifier matches submissions against this catalog; admin-only scenarios (e.g.code_build_strict) deliberately omitexamplesto stay out of it. Add the examples in the same PR that adds the scenario. intent_routes:are first-come. Two scenarios claiming the same route is a hard load error. Thetaskroute is reserved for the tier-3 classifier rule — a plain “task” with no explicit route goes tointent_classifier, which picks by examples; direct callers should pinscenario_idor use a specific route.- Chain depth is capped at 3, and the classifier counts as one hop.
- Non-conversational work types (
rule,fact,planner,metric) bypass the orchestrator entirely.
CI gate
.github/workflows/scenario-replay.yml
runs on every PR touching knowledge-hub/scenarios/**,
knowledge-hub/rubrics/**, or worker/worker/graph/scenarios/**. It loads each
scenario’s golden fixtures from
knowledge-hub/scenarios/<id>/golden/*.json and replays them through
ReplayHarness against CassetteProviders. The job fails on any verifier or
rubric regression beyond declared thresholds.
Impact scope
- Worker: hot path for every paid job goes through
ScenarioOrchestrator. LegacyHierarchyOrchestratoris deleted; the wrappingWorkerOrchestratoris unchanged. - Frontend:
JobDetailPanelmountsScenarioRunTimelinefor every job. Renders nothing on legacy backlog events for backwards compat. - Postgres: one new table
tenant_scenario_overrides(composite PK ontenant_id, scenario_id). - Defaults:
code_buildis the default scenario; override via theAMOS_DEFAULT_SCENARIOenv var or per-jobscenario_idmetadata. - Backwards compat: legacy
node_started/node_finishedevents still fire alongside the newlevel_*events for one release.
Tests
All tests live at repo roottests/. The scenario runtime suite:
test_scenarios.py(28) — schemas, loader, scratchpad.test_scenario_expr.py(25) — gate DSL.test_verifiers.py+test_verifiers_phase2.py(55) — verifier framework + built-ins.test_scenario_compiler.py(10) — compiler topology + end-to-end with mock runners.test_scenario_prompts.py(10) — prompt-builder parity.test_scenario_orchestrator.py(9) — production orchestrator with fake providers.test_scenario_events.py(13) — Phase-4 lifecycle event payloads.test_scenario_phase5.py(12) — fan_out + aggregators + budget kill switch.test_scenario_phase6.py(16) — checksum, hot-reload, tenant overrides, replay harness.test_tenant_overrides.py(12) — DB-backed tenant overrides.test_model_resolver.py(15) — per-level model pinning.test_document_writing.py(15),test_quick_qa_and_research_brief.py(12),test_security_review.py(15) — bundled scenarios end-to-end.
test_provider_registry.py, test_scheduler_producer.py,
test_context_engine.py — all WorkerConfig.__init__ drift not introduced by
the scenario work).