Open-source · MIT · Semantic routing

One endpoint. Every request lands on the right model.

Not the best model overall. The best model for this task. Debug goes to the debugger, prose to the writer — decided per request, from data.

The open-source objection, up front

Running open-source models sounds like a compromise.

Trade a little quality for sovereignty and a zero bill. A good-enough model instead of the frontier one.

That's the assumption CoeOS is built to break.

Not good-enough. Better.

A team of specialists beats a lone generalist.

A single closed model is a generalist. Strong everywhere, exceptional nowhere. It drags its weak axes into every request whether the task needs them or not — and bills for all of it, including the reasoning tokens it burns when the task didn't call for deep thinking.

CoeOS runs the specialist for each skill — the model that scores highest on that exact axis, across 36 open and cloud models on 18 measured disciplines. Debug goes to the debugger. Legal goes to the legal analyst. The request lands once, on the right model, and returns untouched.

Result: 97.2/100 on the TMB benchmark — statistically on par with Fable 5 (96.9) and Claude Opus 4.8 (96.0) on equal-footing API runs. The point-spread is within stochastic noise; what isn't noise is the cost: 4.9× less per test than Fable 5, 1.6× less than Opus 4.8.

Available in two variants. v1.33 (economic) halves the cost by curtailing reasoning-token spend — the hidden line item on "thinking" models. v1.32 (performance) runs heavier, at the same per-test cost as Opus 4.8, but scores above it.

Earned empirically, not claimed

The ranking behind the routing doesn't come from generic leaderboards.

Not MMLU. Not a suffix on a model card. Not a number to screenshot. Dozens of models run through the same battery, scored discipline by discipline, until a real scoreboard emerges — who actually wins where, with the receipt attached. Every routing decision traces back to a score. Not a hunch. A result.

TMB benchmarks · CoeOS v1.3
97.2 score /100 · on par with Fable 5 (96.9) & Opus 4.8 (96.0)
$0.087 per test · vs $0.426 (Fable 5) · $0.141 (Opus 4.8)
42 tests across 5 domains · 18 axes

The TMB benchmarks

More tests than settings — because settling a single axis takes more than a single trial.

42 tests across five domains — agentic execution, code, planning, reasoning, and writing. Each axis has a hand-written rubric, a single run at temperature 0, and an evaluation by Claude Opus 4.8. The result isn't "model X is good" — it's "model X scores 98 on debug and 62 on Swift." Granular enough to route each request to the model proven best at that specific skill, not the highest average.

A pilot, not an aggregator

Here's the line that matters.

Aggregator

Fuses several models per query. Fans the call out, blends the answers, bills for all of it. An algorithm guessing at consensus.

CoeOS pilots.

Reads the request, understands what it is, and steers it — whole, once — to the one model proven best at it. One classification. One call. The answer returns untouched, in the caller's own format.

Semantic routing, not statistical mush. It knows what the request needs before it picks who answers.

Open-source · MIT

Built by the team behind OdyssAI-X.