Skip to main content

Capability first, price second

Each kind of work states a requirement, and the catalog answers with whatever can meet it, cheapest capable first: The order matters and is not a preference: a cheaper model that cannot do the work is not a saving. Requirements are never traded away for price — if only an expensive model can meet them, the work goes there.

Ten providers, one contract

Claude, OpenAI, DeepSeek, GLM (Zhipu), Kimi (Moonshot), MiniMax, plus the Groq and OpenRouter routers. Adding one is a configuration row, not a code change — ten of the eleven speak the same wire format, so a new vendor usually needs no code at all. (Developers: Model routing.) Cost tiers are free → cheap → mid → frontier, and cheap is roughly a tenth of frontier for reasoning-heavy work. That is where the saving comes from: bulk implementation and reasoning prefer cheap, and only the steps that need a frontier model get one. Choosing a tier is choosing a band, not a model. The frontier band holds models from three vendors — deliberately, so the band has somewhere to go when one of them is down (see the fallback section below). Planning and building ride the band’s strongest model; the band’s weaker members are the escape hatch, not a downgrade Amos reaches for willingly.

How hard to think, portably

Every vendor spells reasoning differently — a token budget, an enum, a boolean switch. Amos exposes one ladder, none → low → medium → high → max, and translates. A level asks for high and gets it whether it lands on Claude (a thinking budget), DeepSeek (reasoning_effort), or GLM (a switch that is simply on). Effort is a floor: a model that doesn’t define the requested rung rounds up to the next one it has, because a level asking to think harder should never be answered with less. none is the one exception — it is a ceiling, and never rounds up.

Free means quota, not hardware

“Free” used to mean our own GPU. It now means hosted free tiers, which fail by refusing rather than by billing — so they need pacing, not a budget. A shared Redis ledger tracks requests and tokens per window across every worker, because several workers each keeping their own count would burn a 50-a-day cap in minutes while each believed it had plenty left. A free model whose window is spent is skipped, not tried: the point is to avoid the 429, not to survive it. And a free candidate is admitted on capability before price is considered at all — “it’s free” is never a reason to run work on a model that cannot do it.

Fallback that knows why it failed

Every provider has a health score, and failures decay: one bad afternoon no longer de-prefers a provider forever. Three failures are told apart, because the right response differs:
  • A spent quota skips that model until its window rolls. Not a health failure — the provider is fine, we simply asked too often.
  • An unfunded account (a valid key with a zero balance) says so by name. No amount of retrying fixes it; an operator has to add funds. Vendors report this inconsistently — one sends 402, another sends 429 — so the body is read rather than the status code.
  • A missing capability is refused before the call, naming the model and the feature it lacks, instead of arriving as a provider 400 or a confident answer about an attachment the model never received.
And when the failure is none of those — the provider is healthy, funded and capable, the model is simply unavailable right now — the tier’s own band is the fallback. Dispatch walks its members strongest-first: under the frontier budget a failed request against the strongest model is retried against the band’s next model, then the band’s other vendors, before any provider default or a level failure. The walk crosses providers as it descends, a refused credential skips the rest of that vendor’s candidates in one step, and the step’s UI badge names the model that actually answered. One vendor’s bad afternoon costs a hop instead of the run. A tool loop walks the same band only until its first successful turn; after that the conversation’s tool-call history is vendor-shaped, and switching models mid-conversation would corrupt it. (The shipped bands, the walk’s exact ordering, and its configuration are on Model routing.)

Per-call cost tracking

Every call records provider, model, token counts and cost — plus the effort rung it asked for and the reasoning tokens that rung consumed, so “was thinking harder worth it” is answerable rather than a matter of opinion. Media is recorded by reference and content hash, never as bytes. Prices live in the catalog, which refuses to load a paid model with no price: every OpenAI and Gemini call in this system recorded $0.00 for months because a lookup defaulted to zero, and a cost dashboard that quietly lies is worse than one that fails to start.