Skip to content

TMB Scoreboard

“The cheapest model is the one that doesn’t create rework.” — The Monocle Bear

The TMB Scoreboard (The Monocle Bear) is how we decide which models to actually deploy. Not a leaderboard scraped from a paper — real, demanding tasks graded on strict rubrics, across 18 disciplines, evaluated by Claude Opus 4.8. It is the evidence base behind every routing decision in the stack.

18 axes across five domains:

DomainAxes
Creativecreative writing, rédaction pro
LegalRGPD/conformité, multi-juridictionnel
Codepython, code général, debug, react, swift, refactoring
Planningplan décomposition, plan spec/handoff, jugement/pièges
Agentexécution outillée, discipline/sûreté
+raisonnement/logique, calcul structuré, fast tools

Every axis: score 0–100, single run, strict rubric, graded by Claude Opus 4.8. The result is not “model X is good” but “model X scores 98 on debug, 62 on Swift” — granular enough to route each request to the model proven best at that skill.

ScoreboardAxesWhat it measures
General18Overall ranking — all domains
Coding6Python, code général, debug, React, Swift, refactoring
Planning & Agents7Planning, raisonnement, calcul, agent exec/safety
  1. Kimi-K3-OR — 98.0% avg. Wins or ties on almost every axis. Top creative (98), top legal (99), top planning (98). ⚠ token burner.
  2. Qwen3.8-prev ALI — 97.0%. The strongest local model: runs on-cluster, 97+ on code and planning. ⚠ token burner. 🏠 local.
  3. nemotron3 super OR — 94.8%. Best creative score after Kimi (96), #1 on reasoning (100) and calc (100).
  4. Nex N2 pro OR — 94.8%. #1 agent execution (100). Fast at 134 TPS.
  5. deepseek v4 pro OR — 94.3%. Perfect on debug (100), reasoning (100), calc (100). The most consistent analytical model.

Frontier reference (out of contender ranking): Fable5-OR-REF at 97.7%.

36 contenders spanning local and cloud, including:

  • Kimi — K3 (OR), K2.7 (OR), K2.6 (OR), Linear-48B (local)
  • Qwen — 3.8-prev (ALI, local), 3.7+ / 3.7 Max (OR), 3.5 397B (OR), 3-coder-Next-80B (local), 3.5-122B (local), 3-Next-80B (local)
  • GLM — 5.1 (OR), 5.2 (OR), 5.2 Q6 (local)
  • DeepSeek — v4 pro / v4 flash (OR)
  • Nvidia — nemotron3 super (OR)
  • Nex — N2 pro (OR)
  • MiniMax — minimax3 (MI), minimax2.7 (MI)
  • Lingyi — Ling-3-flash, Ling-2.6 (OR)
  • Other — aion3, mimo2.5, mercury-2, Hy3, step3.7flash, RING2.6, mistral M3.5 / L3, o3, Laguna, destral2, kat-coder-pro-v2.5

Runs on the Odysseus cluster × OdyssAI-X × OpenRouter × direct API. Graded by Claude Opus 4.8 (rejudge pass). Companion memory disabled for every benchmark run.