Skip to main content
Amos’s agents aren’t a flat pool. They’re organized in four levels, each with a defined role.

L1 — Team Lead

Orchestration. Reads the request, classifies it, and decides whether it’s routine or non-trivial (escalates to L2). Delegates work, doesn’t execute it.

L2 — Managerial roles

Leadership. Specializes by role:
  • Architect — system design and structural decisions.
  • Tech Lead — implementation plans and code-level direction.
  • Release Manager — deployment and rollout decisions.
  • QA, Security, and Adversarial leads — quality, risk, and red-team thinking.
L2 doesn’t execute either. It plans and routes work to L4.

L3 — Acceptance

Verification. Reads the executor’s output and checks it against the original ask. Returns pass/fail with deltas (what’s missing, what needs revision). It doesn’t rewrite.

L4 — Executor

Production. Specializes in file writes — emits net-new files and surgical edits — and streams its output as it works, with keep-alive signals so long generations don’t stall.

How escalation works

Within a level, conflicts retry up to three times. After three failures, the conflict moves up — to a more capable agent — rather than burning more attempts at the same level. Sub-job nesting is capped at depth two.

Why this shape

  • Planning and review are short, well-structured tasks. They don’t need the most expensive model, and under a frontier budget they deliberately don’t get one — the budget concentrates where model quality shows.
  • Building is where the model matters. L2 plans and L4 builds; a frontier budget buys frontier quality exactly there.
  • Acceptance is separated from execution. The same agent doesn’t get to grade its own work.
Which model tier serves each level follows the budget you pick (see Free and paid models, together) — the hierarchy is about responsibility, not a fixed cost assignment.

Beyond code

The four-level shape suits code flows. Non-code flows aren’t shaped the same way:
  • Document writing is an editorial chain (outline → research → draft → fact-check → polish).
  • Research briefs fan out across specialist subdomains (web, internal knowledge, academic) and join the results with a synthesis step.
  • Security review is a pipeline: triage, parallel scans, analyst review, human approval, and a final report.
  • Quick Q&A is a single fast level with an automatic format check.
These all run on the same scenario engine. The four-level layering remains the default convention for code-building runs; everything else picks its own arrangement.

Under the hood

The scenario runtime that arranges agent levels per flow is described in Scenarios runtime.