Capability first, price second
Each kind of work states a requirement, and the catalog answers with whatever can meet it, cheapest capable first:
The order matters and is not a preference: a cheaper model that cannot do the
work is not a saving. Requirements are never traded away for price — if only an
expensive model can meet them, the work goes there.
Ten providers, one contract
Claude, OpenAI, DeepSeek, GLM (Zhipu), Kimi (Moonshot), MiniMax, plus the Groq and OpenRouter routers. Adding one is a configuration row, not a code change — ten of the eleven speak the same wire format, so a new vendor usually needs no code at all. (Developers: Model routing.) Cost tiers arefree → cheap → mid → frontier, and cheap is roughly a tenth of
frontier for reasoning-heavy work. That is where the saving comes from: bulk
implementation and reasoning prefer cheap, and only the steps that need a
frontier model get one.
Choosing a tier is choosing a band, not a model. The frontier band holds
models from three vendors — deliberately, so the band has somewhere to go
when one of them is down (see the fallback section below). Planning and
building ride the band’s strongest model; the band’s weaker members are the
escape hatch, not a downgrade Amos reaches for willingly.
How hard to think, portably
Every vendor spells reasoning differently — a token budget, an enum, a boolean switch. Amos exposes one ladder,none → low → medium → high → max, and
translates. A level asks for high and gets it whether it lands on Claude
(a thinking budget), DeepSeek (reasoning_effort), or GLM (a switch that is
simply on).
Effort is a floor: a model that doesn’t define the requested rung rounds
up to the next one it has, because a level asking to think harder should
never be answered with less. none is the one exception — it is a ceiling, and
never rounds up.
Free means quota, not hardware
“Free” used to mean our own GPU. It now means hosted free tiers, which fail by refusing rather than by billing — so they need pacing, not a budget. A shared Redis ledger tracks requests and tokens per window across every worker, because several workers each keeping their own count would burn a 50-a-day cap in minutes while each believed it had plenty left. A free model whose window is spent is skipped, not tried: the point is to avoid the 429, not to survive it. And a free candidate is admitted on capability before price is considered at all — “it’s free” is never a reason to run work on a model that cannot do it.Fallback that knows why it failed
Every provider has a health score, and failures decay: one bad afternoon no longer de-prefers a provider forever. Three failures are told apart, because the right response differs:- A spent quota skips that model until its window rolls. Not a health failure — the provider is fine, we simply asked too often.
- An unfunded account (a valid key with a zero balance) says so by name. No amount of retrying fixes it; an operator has to add funds. Vendors report this inconsistently — one sends 402, another sends 429 — so the body is read rather than the status code.
- A missing capability is refused before the call, naming the model and the feature it lacks, instead of arriving as a provider 400 or a confident answer about an attachment the model never received.