Fuses several models per query. Fans the call out, blends the answers, bills for all of it. An algorithm guessing at consensus.
Open-source · MIT · Semantic routing
Not the best model overall. The best model for this task. Debug goes to the debugger, prose to the writer — decided per request, from data.
The open-source objection, up front
Trade a little quality for sovereignty and a zero bill. A good-enough model instead of the frontier one.
That's the assumption CoeOS is built to break.
Not good-enough. Better.
A single closed model is a generalist. Strong everywhere, exceptional nowhere. It drags its weak axes into every request whether the task needs them or not — and bills for all of it, including the reasoning tokens it burns when the task didn't call for deep thinking.
CoeOS runs the specialist for each skill — the model that scores highest on that exact axis, across 36 open and cloud models on 18 measured disciplines. Debug goes to the debugger. Legal goes to the legal analyst. The request lands once, on the right model, and returns untouched.
Result: 97.2/100 on the TMB benchmark — statistically on par with Fable 5 (96.9) and Claude Opus 4.8 (96.0) on equal-footing API runs. The point-spread is within stochastic noise; what isn't noise is the cost: 4.9× less per test than Fable 5, 1.6× less than Opus 4.8.
Available in two variants. v1.33 (economic) halves the cost by curtailing reasoning-token spend — the hidden line item on "thinking" models. v1.32 (performance) runs heavier, at the same per-test cost as Opus 4.8, but scores above it.
Earned empirically, not claimed
Not MMLU. Not a suffix on a model card. Not a number to screenshot. Dozens of models run through the same battery, scored discipline by discipline, until a real scoreboard emerges — who actually wins where, with the receipt attached. Every routing decision traces back to a score. Not a hunch. A result.
The TMB benchmarks
42 tests across five domains — agentic execution, code, planning, reasoning, and writing. Each axis has a hand-written rubric, a single run at temperature 0, and an evaluation by Claude Opus 4.8. The result isn't "model X is good" — it's "model X scores 98 on debug and 62 on Swift." Granular enough to route each request to the model proven best at that specific skill, not the highest average.
A pilot, not an aggregator
Fuses several models per query. Fans the call out, blends the answers, bills for all of it. An algorithm guessing at consensus.
Reads the request, understands what it is, and steers it — whole, once — to the one model proven best at it. One classification. One call. The answer returns untouched, in the caller's own format.
Semantic routing, not statistical mush. It knows what the request needs before it picks who answers.
Open-source · MIT