Skip to main content

Which flow runs

Amos picks the flow for you, in order of precedence:
  1. Explicit choice — a pinned scenario id (used by replays and admin tools).
  2. What you asked for — the task type from the composer maps to a flow: a project build → code_build, a document → document_writing, a question → quick_qa, research → research_brief, a security review → security_review.
  3. The default — a code build.
The run’s timeline records why a flow was picked, so “why did it do that?” is answerable. One flow can also chain into another mid-run — an intent classification flowing into a question-answering flow, say — and the timeline draws the link.

The steps

A flow is a graph of steps (levels), not a straight line. Each step has a role on the team — a lead that frames the work, an architect that plans, a builder that writes, a reviewer that judges — and a declared kind:
  • Prompt steps produce written output: the brief, the plan, the review.
  • Build steps are where the model actually works: an agentic tool loop that reads, writes, and edits files in the project workspace. The filesystem is the artifact — downstream steps receive the touched paths, not file dumps.
  • Verifier steps check the result without a model: does it build, does it typecheck, is every route reachable. A failure feeds the specifics back to the builder for a fix, with a bounded retry budget before the conflict escalates upward.
  • Gate steps decide where the run goes next — a verdict of PASS routes onward, FAIL routes back. Every branch taken is recorded as a gate decision in the timeline.
  • Fan-out steps run children in parallel and merge them (majority pass, weighted score, all pass).
  • Approval steps stop the run and ask a human.

What you see while it runs

The run canvas and step list tell the whole story live, without polling:
  • Each step’s node wears the model’s brand mark, and the step row names the model that ran it. If dispatch had to fall to the next model in the band, the step corrects itself and badges the model that actually answered.
  • Steps the worker composed in code — deliberately, with no model call — are labelled “composed in code” rather than left looking like a missing mark.
  • Verifier rows show pass/fail with scores and reasons; gate rows show which branch fired and why; the budget ticker accumulates cost, tokens, and time.

Confirmations and approvals

A run parks — and tells you — whenever a human decision is needed:
  • Before the work starts: confirming intent, picking a template and design, and confirming the feature list are all user confirmations.
  • Mid-run: scope checks, clarification questions, and approval steps all park the run. Every parked gate becomes a durable, addressable approval — so the decision can land from the run’s own panel, the approvals page, or Slack, not only from the live stream. Answering resumes the run exactly where it paused.

Budgets, stopping, and cancellation

Every scenario declares caps on cost, tokens, and duration. Hitting a cap aborts the run — remaining steps skip, and the run ends as aborted with the reason. You can also press Stop at any moment: the run ends as cancelled and partial work is discarded (nothing is committed or pushed) — only a fresh prompt starts over.

When it finishes

The run ends completed, and the timeline keeps the full story: which model ran each step, which verifiers passed or failed with what scores, which branches the gates took, what files were created, where it was deployed. Archived runs preserve the same per-step detail, so a run from last week reads as clearly as one from a minute ago.

Under the hood