TMB Scoreboard — Coding
The Monocle Bear · updated 29 July 2026
- Axes: python, code_general, debug, react, swift, refactoring (6 axes, /100 each)
- Evaluator: Claude Opus 4.8 (rejudge pass)
- Scale: 0–100 per axis
Code scoreboard
Section titled “Code scoreboard”| # | Model | Avg code | Python | Code gen. | Debug | React | Swift | Refact. | TPS |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Kimi-K3-OR | 98.0 | 98 | 98 | 98 | 98 | 99 | 97 | 31.2 |
| 2 | Qwen3.8-prev ALI 🏠 | 97.5 | 97 | 98 | 100 | 97 | 96 | 97 | 44.8 |
| 3 | glm-5.1-OR | 95.4 | 96 | 94 | 100 | 91 | 93 | 98 | 52.4 |
| 4 | kimi-k2.6-OR | 94.5 | 97 | 95 | 93 | 94 | 92 | 96 | 42.7 |
| 5 | Nex N2 pro OR | 94.3 | 96 | 94 | 96 | 98 | 88 | — | 134.1 |
| 6 | deepseek v4 pro OR | 93.9 | 95 | 91 | 100 | 98 | 84 | 96 | 52.2 |
| 7 | GLM5.2-OR | 93.9 | 97 | 92 | 98 | 94 | 88 | 95 | 41.9 |
| 8 | nemotron3 super OR | 92.6 | 96 | 92 | 96 | 90 | 88 | 94 | 112.5 |
| 9 | KIMI 2.7 OR | 92.5 | 92 | 92 | 98 | 90 | 90 | 93 | 48.4 |
| 10 | Qwen3.7+ OR | 92.2 | 95 | 91 | 92 | 98 | 85 | 92 | 54.1 |
| 11 | aion3 OR | 91.9 | 93 | 91 | 100 | 93 | 86 | 88 | 46.4 |
| 12 | Qwen3.7Max OR | 90.3 | 95 | 89 | 88 | 92 | 93 | 85 | 54.0 |
| 13 | deepseek v4 flash OR | 88.2 | 94 | 88 | 86 | 87 | 82 | 92 | 58.2 |
| 14 | minimax3 MI | 87.9 | 95 | 86 | 94 | 94 | 58 | 100 | 68.1 |
| 15 | Ling-3-flash - OR | 85.8 | 94 | 83 | 100 | 90 | 62 | 86 | 298.9 |
| 16 | mimo2.5 XI | 85.7 | 96 | 86 | 93 | 96 | 45 | 98 | 75.0 |
| 17 | minimax2.7 MI | 84.9 | 96 | 86 | 81 | 90 | 85 | 72 | 48.4 |
| 18 | GLM5.2 Q6 🏠 | 84.0 | 90 | 82 | 86 | 86 | 75 | 85 | 11.2 |
| 19 | mercury-2-OR | 83.9 | 84 | 82 | 94 | 74 | 69 | 100 | 433.9 |
| 20 | step3.7flash OR | 83.5 | 88 | 84 | 82 | 86 | 80 | 82 | 170.9 |
| 21 | qwen3.5 397 OR | 83.2 | 75 | 81 | 100 | 90 | 74 | 80 | 50.9 |
| 22 | Ling-2.6 - OR | 81.6 | 94 | 83 | 89 | 84 | 76 | 64 | 108.3 |
| 23 | laguna-xs-2.1 🏠 | 81.4 | 93 | 82 | 78 | 88 | 67 | 80 | 79.4 |
| 24 | Hy3 - OR | 81.1 | 95 | 81 | 78 | 80 | 65 | 87 | 51.4 |
| 25 | Qwen3-coder-Next-80b 🏠 | 79.6 | 88 | 81 | 77 | 89 | 62 | 80 | 58.6 |
| 26 | mistral M3.5 | 79.4 | 83 | 78 | 72 | 88 | 65 | 90 | 144.9 |
| 27 | mistral L3 OR | 77.1 | 93 | 76 | 84 | 86 | 71 | 52 | 64.8 |
| 28 | o3-OR | 76.7 | 94 | 76 | 92 | 88 | 35 | 76 | 132.8 |
| 29 | Qwen3-Next-80B-A3B-Instruct-8bit 🏠 | 76.5 | 91 | 78 | 80 | 82 | 68 | 60 | 61.3 |
| 30 | Qwen3.5-122B-8H16 🏠 | 75.4 | 94 | 76 | 74 | 78 | 60 | 70 | 41.2 |
| 31 | Laguna-S2.1 🏠 | 74.5 | 83 | 71 | 84 | 75 | 50 | 83 | 42.4 |
| 32 | RING2.6 OR | 73.1 | 89 | 77 | 88 | 32 | 70 | 83 | 125.5 |
| 33 | lagunaM1 OR | 69.7 | 80 | 70 | 81 | 72 | 43 | 72 | 42.2 |
| 34 | Kimi-Linear-48B 🏠 | 65.2 | 80 | 68 | 64 | 41 | 63 | 75 | 82.1 |
| 35 | destral2 OR | 65.2 | 86 | 68 | 56 | 78 | 54 | 50 | 71.3 |
| 36 | kat-coder-pro-v2.5-OR | 63.0 | 92 | 55 | 91 | 24 | 36 | 80 | 109.1 |
References
Section titled “References”Via OpenRouter — directly comparable with contenders.
| Model | Avg code | Python | Code gen. | Debug | React | Swift | Refact. |
|---|---|---|---|---|---|---|---|
| Fusion-REF | 97.7 | 99 | 98 | 100 | 96 | 96 | 98 |
| GPT 5.6 Sol-REF | 97.4 | 97 | 97 | 98 | 98 | 98 | 96 |
| Fable5-OR-REF | 96.7 | 97 | 97 | 98 | 97 | 96 | 96 |
| GPT 5.5-REF | 96.6 | 97 | 96 | 98 | 98 | 95 | 96 |
| Opus5-OR-REF | 93.0 | 98 | 93 | 99 | 96 | 72 | 99 |
| Opus-4.8-OR-REF | 92.8 | 97 | 93 | 92 | 93 | 88 | 93 |
Run under Claude Code harness (not directly comparable)
Section titled “Run under Claude Code harness (not directly comparable)”These runs went through the Claude Code harness rather than direct API — not on the same footing as the contenders above.
| Model | Avg code | Python | Code gen. | Debug | React | Swift | Refact. |
|---|---|---|---|---|---|---|---|
| Fable5-REF | 98.2 | 99 | 98 | 100 | 96 | 97 | 99 |
| Opus4.8-REF | 97.4 | 97 | 97 | 100 | 98 | 96 | 96 |
TMB Benchmark v5 — The Monocle Bear — July 2026.