TMB Scoreboard — General
The Monocle Bear · updated 29 July 2026
- Axes: 18 disciplines — creative, legal, code (6), planning (3), agent (2), reasoning, calc, tools
- Evaluator: Claude Opus 4.8 (rejudge pass)
- Scale: 0–100 per axis, arithmetic mean (verified axes only)
- Infrastructure: Odysseus cluster × OdyssAI-X × OpenRouter × direct API
Main scoreboard — all contenders (≥ 16 axes)
Section titled “Main scoreboard — all contenders (≥ 16 axes)”Avg = arithmetic mean of verified axes. ⚠ = token burner. 🏠 = local model.
Frontier reference ceiling: Fable5-OR-REF at 97.7% (out of contender ranking — see below).
| # | Model | Avg | Creative | Legal | Code | Planning | Agent | TPS | |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Kimi-K3-OR | 98.0 | 98 | 99 | 98 | 98 | 99 | 31.2 | ⚠ |
| 2 | Qwen3.8-prev ALI | 97.0 | 95 | 96 | 98 | 98 | 95 | 44.8 | ⚠ 🏠 |
| 3 | nemotron3 super OR | 94.8 | 96 | 97 | 93 | 93 | 96 | 112.5 | |
| 4 | Nex N2 pro OR | 94.8 | 94 | 92 | 94 | 94 | 100 | 134.1 | ⚠ |
| 5 | deepseek v4 pro OR | 94.3 | 94 | 93 | 94 | 91 | 97 | 52.2 | |
| 6 | glm-5.1-OR | 93.5 | 84 | 93 | 95 | 89 | 98 | 52.4 | ⚠ |
| 7 | KIMI 2.7 OR | 93.1 | 90 | 95 | 92 | 92 | 97 | 48.4 | |
| 8 | GLM5.2-OR | 91.7 | 91 | 84 | 94 | 83 | 98 | 41.9 | |
| 9 | kimi-k2.6-OR | 91.7 | 73 | 90 | 94 | 91 | 97 | 42.7 | |
| 10 | Qwen3.7Max OR | 90.4 | 88 | 86 | 90 | 85 | 98 | 54.0 | |
| 11 | Qwen3.7+ OR | 90.1 | 80 | 81 | 92 | 89 | 98 | 54.1 | |
| 12 | qwen3.5 397 OR | 89.3 | 90 | 88 | 83 | 89 | 98 | 50.9 | |
| 13 | Ling-3-flash - OR | 89.2 | 78 | 94 | 86 | 92 | 95 | 298.9 | ⚠ |
| 14 | minimax3 MI | 88.6 | 89 | 93 | 88 | 94 | 94 | 68.1 | |
| 15 | aion3 OR | 87.6 | 80 | 76 | 92 | 81 | 94 | 46.4 | |
| 16 | GLM5.2 Q6 | 87.6 | 80 | 86 | 84 | 87 | 97 | 11.2 | 🏠 |
| 17 | step3.7flash OR | 87.3 | 76 | 90 | 84 | 90 | 95 | 170.9 | |
| 18 | RING2.6 OR | 87.1 | 89 | 98 | 73 | 93 | 97 | 125.5 | ⚠ |
| 19 | minimax2.7 MI | 86.5 | 67 | 92 | 85 | 94 | 87 | 48.4 | |
| 20 | mimo2.5 XI | 85.8 | 60 | 100 | 86 | 93 | 97 | 75.0 | ⚠ |
| 21 | Hy3 - OR | 85.3 | 87 | 88 | 81 | 87 | 96 | 51.4 | |
| 22 | Ling-2.6 - OR | 85.1 | 89 | 88 | 82 | 94 | 96 | 108.3 | |
| 23 | deepseek v4 flash OR | 85.1 | 89 | 89 | 88 | 77 | 98 | 58.2 | |
| 24 | o3-OR | 84.0 | 58 | 89 | 77 | 98 | 98 | 132.8 | |
| 25 | Qwen3-coder-Next-80b | 82.8 | 86 | 92 | 80 | 86 | 93 | 58.6 | 🏠 |
| 26 | mercury-2-OR | 82.5 | 67 | 84 | 84 | 88 | 95 | 433.9 | |
| 27 | Qwen3-Next-80B-A3B-Instruct-8bit | 81.9 | 82 | 87 | 76 | 89 | 98 | 61.3 | 🏠 |
| 28 | mistral M3.5 | 81.6 | 85 | 93 | 79 | 81 | 93 | 144.9 | |
| 29 | mistral L3 OR | 80.8 | 80 | 88 | 77 | 82 | 96 | 64.8 | |
| 30 | laguna-xs-2.1 | 79.6 | 67 | 82 | 81 | 85 | 90 | 79.4 | 🏠 |
| 31 | kat-coder-pro-v2.5-OR | 79.4 | 83 | 94 | 63 | 88 | 98 | 109.1 | |
| 32 | Qwen3.5-122B-8H16 | 77.8 | 58 | 89 | 75 | 89 | 84 | 41.2 | 🏠 |
| 33 | lagunaM1 OR | 76.8 | 86 | 80 | 70 | 81 | 99 | 42.2 | |
| 34 | destral2 OR | 75.6 | 76 | 86 | 65 | 84 | 96 | 71.3 | |
| 35 | Laguna-S2.1 | 73.2 | 76 | 67 | 75 | 80 | 90 | 42.4 | 🏠 |
| 36 | Kimi-Linear-48B | 70.7 | 71 | 64 | 65 | 78 | 90 | 82.1 | 🏠 |
References (out of contender ranking)
Section titled “References (out of contender ranking)”Via OpenRouter — directly comparable with contenders.
| Model | Avg | Creative | Legal | Code | Planning | Agent |
|---|---|---|---|---|---|---|
| Fable5-OR-REF | 97.7 | 99 | 97 | 97 | 100 | 95 |
| GPT 5.6 Sol-REF | 96.3 | 94 | 97 | 97 | 98 | 97 |
| Opus-4.8-OR-REF | 96.2 | 100 | 96 | 93 | 99 | 96 |
| GPT 5.5-REF | 95.5 | 94 | 97 | 97 | 95 | 97 |
| Fusion-REF | 94.6 | 85 | 98 | 98 | 94 | 91 |
| Opus5-OR-REF | 94.2 | 95 | 96 | 93 | 94 | 97 |
Run under Claude Code harness (not directly comparable)
Section titled “Run under Claude Code harness (not directly comparable)”These runs went through the Claude Code harness rather than direct API — not on the same footing as the contenders above.
| Model | Avg | Creative | Legal | Code | Planning | Agent |
|---|---|---|---|---|---|---|
| Fable5-REF | 97.9 | 98 | 99 | 98 | 98 | 98 |
| Opus4.8-REF | 97.5 | 100 | 96 | 97 | 97 | 99 |
Per-axis champions
Section titled “Per-axis champions”| Axis | #1 | #2 | #3 |
|---|---|---|---|
| Créatif / fiction longue | Kimi-K3-OR (98) | nemotron3 super OR (92) | Qwen3.8-prev ALI (90) |
| Rédaction pro / business | Qwen3.8-prev ALI (100) | RING2.6 OR (100) | deepseek v4 pro OR (100) |
| RGPD / conformité | Kimi-K3-OR (100) | mimo2.5 XI (100) | RING2.6 OR (98) |
| Conformité multi-jur. | RING2.6 OR (99) | mimo2.5 XI (99) | Kimi-K3-OR (98) |
| Raisonnement / logique | GLM5.2 Q6 (100) | Qwen3.8-prev ALI (100) | deepseek v4 pro OR (100) |
| Calcul / structuré | Qwen3.5-122B-8H16 (100) | Qwen3.7+ OR (100) | Qwen3.7Max OR (100) |
| Python / scripts | Kimi-K3-OR (98) | kimi-k2.6-OR (97) | Qwen3.8-prev ALI (97) |
| Code général | Kimi-K3-OR (98) | Qwen3.8-prev ALI (98) | kimi-k2.6-OR (95) |
| Debug | Ling-3-flash - OR (100) | Qwen3.8-prev ALI (100) | deepseek v4 pro OR (100) |
| React / front | Kimi-K3-OR (98) | Qwen3.7+ OR (98) | deepseek v4 pro OR (98) |
| Swift | Kimi-K3-OR (99) | Qwen3.8-prev ALI (96) | Qwen3.7Max OR (93) |
| Refactoring | mercury-2-OR (100) | minimax3 MI (100) | glm-5.1-OR (98) |
| Planning — décomposition | KIMI 2.7 OR (100) | Kimi-K3-OR (100) | Ling-2.6 - OR (98) |
| Planning — handoff / spec | o3-OR (100) | Qwen3-coder-Next-80b (100) | Ling-2.6 - OR (98) |
| Planning — jugement / pièges | Kimi-K3-OR (100) | Qwen3.8-prev ALI (99) | minimax2.7 MI (97) |
| Agent — exécution outillée | kimi-k2.6-OR (100) | Nex N2 pro OR (100) | kat-coder-pro-v2.5-OR (99) |
| Agent — discipline / sûreté | GLM5.2 Q6 (100) | Hy3 - OR (100) | Kimi-K3-OR (100) |
TMB Benchmark v5 — The Monocle Bear — July 2026. 18 axes · 36 contenders · Evaluator: Claude Opus 4.8.