Which flow runs
Amos picks the flow for you, in order of precedence:- Explicit choice — a pinned scenario id (used by replays and admin tools).
- What you asked for — the task type from the composer maps to a flow:
a project build →
code_build, a document →document_writing, a question →quick_qa, research →research_brief, a security review →security_review. - The default — a code build.
The steps
A flow is a graph of steps (levels), not a straight line. Each step has a role on the team — a lead that frames the work, an architect that plans, a builder that writes, a reviewer that judges — and a declared kind:- Prompt steps produce written output: the brief, the plan, the review.
- Build steps are where the model actually works: an agentic tool loop that reads, writes, and edits files in the project workspace. The filesystem is the artifact — downstream steps receive the touched paths, not file dumps.
- Verifier steps check the result without a model: does it build, does it typecheck, is every route reachable. A failure feeds the specifics back to the builder for a fix, with a bounded retry budget before the conflict escalates upward.
- Gate steps decide where the run goes next — a verdict of PASS routes onward, FAIL routes back. Every branch taken is recorded as a gate decision in the timeline.
- Fan-out steps run children in parallel and merge them (majority pass, weighted score, all pass).
- Approval steps stop the run and ask a human.
What you see while it runs
The run canvas and step list tell the whole story live, without polling:- Each step’s node wears the model’s brand mark, and the step row names the model that ran it. If dispatch had to fall to the next model in the band, the step corrects itself and badges the model that actually answered.
- Steps the worker composed in code — deliberately, with no model call — are labelled “composed in code” rather than left looking like a missing mark.
- Verifier rows show pass/fail with scores and reasons; gate rows show which branch fired and why; the budget ticker accumulates cost, tokens, and time.
Confirmations and approvals
A run parks — and tells you — whenever a human decision is needed:- Before the work starts: confirming intent, picking a template and design, and confirming the feature list are all user confirmations.
- Mid-run: scope checks, clarification questions, and approval steps all park the run. Every parked gate becomes a durable, addressable approval — so the decision can land from the run’s own panel, the approvals page, or Slack, not only from the live stream. Answering resumes the run exactly where it paused.
Budgets, stopping, and cancellation
Every scenario declares caps on cost, tokens, and duration. Hitting a cap aborts the run — remaining steps skip, and the run ends asaborted with the
reason. You can also press Stop at any moment: the run ends as
cancelled and partial work is discarded (nothing is committed or pushed) —
only a fresh prompt starts over.
When it finishes
The run endscompleted, and the timeline keeps the full story: which model
ran each step, which verifiers passed or failed with what scores, which
branches the gates took, what files were created, where it was deployed.
Archived runs preserve the same per-step detail, so a run from last week reads
as clearly as one from a minute ago.
Under the hood
- Scenarios runtime — the engine: flows as YAML, compiled and executed per job; verifiers; replay in CI.
- HTTP API & event stream — the event contract the canvas renders.
- Model routing — how each step’s model and its fallback band are chosen.