Skip to content

TMB Scoreboard — General

The Monocle Bear · updated 29 July 2026

  • Axes: 18 disciplines — creative, legal, code (6), planning (3), agent (2), reasoning, calc, tools
  • Evaluator: Claude Opus 4.8 (rejudge pass)
  • Scale: 0–100 per axis, arithmetic mean (verified axes only)
  • Infrastructure: Odysseus cluster × OdyssAI-X × OpenRouter × direct API

Main scoreboard — all contenders (≥ 16 axes)

Section titled “Main scoreboard — all contenders (≥ 16 axes)”

Avg = arithmetic mean of verified axes. ⚠ = token burner. 🏠 = local model.

Frontier reference ceiling: Fable5-OR-REF at 97.7% (out of contender ranking — see below).

#ModelAvgCreativeLegalCodePlanningAgentTPS
1Kimi-K3-OR98.0989998989931.2
2Qwen3.8-prev ALI97.0959698989544.8⚠ 🏠
3nemotron3 super OR94.89697939396112.5
4Nex N2 pro OR94.894929494100134.1
5deepseek v4 pro OR94.3949394919752.2
6glm-5.1-OR93.5849395899852.4
7KIMI 2.7 OR93.1909592929748.4
8GLM5.2-OR91.7918494839841.9
9kimi-k2.6-OR91.7739094919742.7
10Qwen3.7Max OR90.4888690859854.0
11Qwen3.7+ OR90.1808192899854.1
12qwen3.5 397 OR89.3908883899850.9
13Ling-3-flash - OR89.27894869295298.9
14minimax3 MI88.6899388949468.1
15aion3 OR87.6807692819446.4
16GLM5.2 Q687.6808684879711.2🏠
17step3.7flash OR87.37690849095170.9
18RING2.6 OR87.18998739397125.5
19minimax2.7 MI86.5679285948748.4
20mimo2.5 XI85.86010086939775.0
21Hy3 - OR85.3878881879651.4
22Ling-2.6 - OR85.18988829496108.3
23deepseek v4 flash OR85.1898988779858.2
24o3-OR84.05889779898132.8
25Qwen3-coder-Next-80b82.8869280869358.6🏠
26mercury-2-OR82.56784848895433.9
27Qwen3-Next-80B-A3B-Instruct-8bit81.9828776899861.3🏠
28mistral M3.581.68593798193144.9
29mistral L3 OR80.8808877829664.8
30laguna-xs-2.179.6678281859079.4🏠
31kat-coder-pro-v2.5-OR79.48394638898109.1
32Qwen3.5-122B-8H1677.8588975898441.2🏠
33lagunaM1 OR76.8868070819942.2
34destral2 OR75.6768665849671.3
35Laguna-S2.173.2766775809042.4🏠
36Kimi-Linear-48B70.7716465789082.1🏠

Via OpenRouter — directly comparable with contenders.

ModelAvgCreativeLegalCodePlanningAgent
Fable5-OR-REF97.799979710095
GPT 5.6 Sol-REF96.39497979897
Opus-4.8-OR-REF96.210096939996
GPT 5.5-REF95.59497979597
Fusion-REF94.68598989491
Opus5-OR-REF94.29596939497

Run under Claude Code harness (not directly comparable)

Section titled “Run under Claude Code harness (not directly comparable)”

These runs went through the Claude Code harness rather than direct API — not on the same footing as the contenders above.

ModelAvgCreativeLegalCodePlanningAgent
Fable5-REF97.99899989898
Opus4.8-REF97.510096979799

Axis#1#2#3
Créatif / fiction longueKimi-K3-OR (98)nemotron3 super OR (92)Qwen3.8-prev ALI (90)
Rédaction pro / businessQwen3.8-prev ALI (100)RING2.6 OR (100)deepseek v4 pro OR (100)
RGPD / conformitéKimi-K3-OR (100)mimo2.5 XI (100)RING2.6 OR (98)
Conformité multi-jur.RING2.6 OR (99)mimo2.5 XI (99)Kimi-K3-OR (98)
Raisonnement / logiqueGLM5.2 Q6 (100)Qwen3.8-prev ALI (100)deepseek v4 pro OR (100)
Calcul / structuréQwen3.5-122B-8H16 (100)Qwen3.7+ OR (100)Qwen3.7Max OR (100)
Python / scriptsKimi-K3-OR (98)kimi-k2.6-OR (97)Qwen3.8-prev ALI (97)
Code généralKimi-K3-OR (98)Qwen3.8-prev ALI (98)kimi-k2.6-OR (95)
DebugLing-3-flash - OR (100)Qwen3.8-prev ALI (100)deepseek v4 pro OR (100)
React / frontKimi-K3-OR (98)Qwen3.7+ OR (98)deepseek v4 pro OR (98)
SwiftKimi-K3-OR (99)Qwen3.8-prev ALI (96)Qwen3.7Max OR (93)
Refactoringmercury-2-OR (100)minimax3 MI (100)glm-5.1-OR (98)
Planning — décompositionKIMI 2.7 OR (100)Kimi-K3-OR (100)Ling-2.6 - OR (98)
Planning — handoff / speco3-OR (100)Qwen3-coder-Next-80b (100)Ling-2.6 - OR (98)
Planning — jugement / piègesKimi-K3-OR (100)Qwen3.8-prev ALI (99)minimax2.7 MI (97)
Agent — exécution outilléekimi-k2.6-OR (100)Nex N2 pro OR (100)kat-coder-pro-v2.5-OR (99)
Agent — discipline / sûretéGLM5.2 Q6 (100)Hy3 - OR (100)Kimi-K3-OR (100)

TMB Benchmark v5 — The Monocle Bear — July 2026. 18 axes · 36 contenders · Evaluator: Claude Opus 4.8.