TMB Scoreboard
“The cheapest model is the one that doesn’t create rework.” — The Monocle Bear
The TMB Scoreboard (The Monocle Bear) is how we decide which models to actually deploy. Not a leaderboard scraped from a paper — real, demanding tasks graded on strict rubrics, across 18 disciplines, evaluated by Claude Opus 4.8. It is the evidence base behind every routing decision in the stack.
What we measure
Section titled “What we measure”18 axes across five domains:
| Domain | Axes |
|---|---|
| Creative | creative writing, rédaction pro |
| Legal | RGPD/conformité, multi-juridictionnel |
| Code | python, code général, debug, react, swift, refactoring |
| Planning | plan décomposition, plan spec/handoff, jugement/pièges |
| Agent | exécution outillée, discipline/sûreté |
| + | raisonnement/logique, calcul structuré, fast tools |
Every axis: score 0–100, single run, strict rubric, graded by Claude Opus 4.8. The result is not “model X is good” but “model X scores 98 on debug, 62 on Swift” — granular enough to route each request to the model proven best at that skill.
The three scoreboards
Section titled “The three scoreboards”| Scoreboard | Axes | What it measures |
|---|---|---|
| General | 18 | Overall ranking — all domains |
| Coding | 6 | Python, code général, debug, React, Swift, refactoring |
| Planning & Agents | 7 | Planning, raisonnement, calcul, agent exec/safety |
Top 5 — July 2026
Section titled “Top 5 — July 2026”- Kimi-K3-OR — 98.0% avg. Wins or ties on almost every axis. Top creative (98), top legal (99), top planning (98). ⚠ token burner.
- Qwen3.8-prev ALI — 97.0%. The strongest local model: runs on-cluster, 97+ on code and planning. ⚠ token burner. 🏠 local.
- nemotron3 super OR — 94.8%. Best creative score after Kimi (96), #1 on reasoning (100) and calc (100).
- Nex N2 pro OR — 94.8%. #1 agent execution (100). Fast at 134 TPS.
- deepseek v4 pro OR — 94.3%. Perfect on debug (100), reasoning (100), calc (100). The most consistent analytical model.
Frontier reference (out of contender ranking): Fable5-OR-REF at 97.7%.
Models tested
Section titled “Models tested”36 contenders spanning local and cloud, including:
- Kimi — K3 (OR), K2.7 (OR), K2.6 (OR), Linear-48B (local)
- Qwen — 3.8-prev (ALI, local), 3.7+ / 3.7 Max (OR), 3.5 397B (OR), 3-coder-Next-80B (local), 3.5-122B (local), 3-Next-80B (local)
- GLM — 5.1 (OR), 5.2 (OR), 5.2 Q6 (local)
- DeepSeek — v4 pro / v4 flash (OR)
- Nvidia — nemotron3 super (OR)
- Nex — N2 pro (OR)
- MiniMax — minimax3 (MI), minimax2.7 (MI)
- Lingyi — Ling-3-flash, Ling-2.6 (OR)
- Other — aion3, mimo2.5, mercury-2, Hy3, step3.7flash, RING2.6, mistral M3.5 / L3, o3, Laguna, destral2, kat-coder-pro-v2.5
Infrastructure & evaluators
Section titled “Infrastructure & evaluators”Runs on the Odysseus cluster × OdyssAI-X × OpenRouter × direct API. Graded by Claude Opus 4.8 (rejudge pass). Companion memory disabled for every benchmark run.
Read next
Section titled “Read next”- General scoreboard →
- Coding scoreboard →
- Planning & Agents scoreboard →
- CoeOS → — how these scores become a virtual model.