What a scenario can express
- An ordered sequence of levels, each with a plain-language name and a one-line description shown in the UI.
- Branching on typed conditions — verdicts, confidence thresholds.
- Parallel fan-out to child levels, joined by an aggregator.
- A human-approval checkpoint with an approver role.
- Per-level model pinning with a fallback band.
- Quality gates: deterministic checks (build, typecheck, output format) run serially, in parallel, or as a race; an LLM judge can score a level against a versioned rubric, blocking or advisory.
- A budget kill switch (cost, tokens, duration) and a per-level failure policy: retry, escalate, continue, or abort.
What you get for free
- Live run detail — every level and every check reports progress as per-level cards in the job detail panel.
- Live cost + token telemetry — the budget chip updates live, flips red when the budget is exceeded, and the run aborts.
- Reproducibility — each run records the exact scenario version and checksum that produced it.
- Tenant overrides — budgets, constraints, and verifier weights can be tuned per customer without forking the base scenario.
How the run canvas reads a scenario
The canvas lays the flow out as an actor-lane grid — one column per role tier (L1, L2, …, in first-appearance order), the decision graph in place before the first level runs. Three touches keep a live run legible:- Named steps. Cards and timeline entries read “TEAM LEAD · Write the brief” instead of a second indistinguishable “Team Lead”; the description is the card’s hover tooltip and toggles that step’s live output in the timeline.
- Dead branches lift out of the way. Levels the run never visited sit on the band above the spine instead of between live steps — resolved at run end, with the nodes gliding rather than jumping.
- Only ambiguous edges are arrowed. Forward flow — left to right across lanes, top to bottom within one — carries no arrowhead; the retries and escalations that double back do, tinted to their live state.
What’s bundled
Seven scenarios cover code, writing, research, security, and Q&A (the three rubrics are versioned separately):code_build— brief → plan → execute → build verify → accept. The default for paid project jobs.code_build_strict— adds parallel build + typecheck + lint and an LLM-judge architecture-review signal.document_writing— outline → research → draft → fact check → polish.research_brief— decompose → fan out across web, knowledge base, and academic sources → synthesis with a rigor signal.quick_qa— one free-tier level with an output-format check; sub-second answers.security_review— triage → parallel scanner passes → Opus-pinned analysis → human approval gate → final report.parallel_review_demo— a fan-out test scaffold, not a production flow.