Skip to content

TMB Scoreboard — Planning & Agents

The Monocle Bear · updated 29 July 2026

  • Axes: plan_decompo, plan_spec, plan_judgment, reasoning, calc, agent_exec, agent_safety (7 axes)
  • Evaluator: Claude Opus 4.8 (rejudge pass)
  • Scale: 0–100 per axis

#ModelAvgP.decompoP.specP.judgmentReasoningCalcAgent execAgent safeTPS
1Kimi-K3-OR97.71009410094989810031.2
2Qwen3.8-prev ALI 🏠97.398969910098969444.8
3o3-OR96.3981009590969898132.8
4Nex N2 pro OR96.2969888969610099134.1
5nemotron3 super OR95.89694891001009299112.5
6RING2.6 OR95.392949494989698125.5
7kimi-k2.6-OR95.3849892981001009542.7
8deepseek v4 pro OR95.1889886100100959952.2
9qwen3.5 397 OR94.78494891001009610050.9
10glm-5.1-OR94.68296881001009610052.4
11KIMI 2.7 OR94.210094829496969848.4
12Ling-3-flash - OR94.092968794999199298.9
13step3.7flash OR94.09685949891100170.9
14Qwen3.7+ OR93.778929694100989854.1
15GLM5.2 Q6 🏠93.1869678100989410011.2
16Qwen3.7Max OR92.378948494100989954.0
17GLM5.2-OR91.9789278100100969941.9
18minimax2.7 MI91.99094979494789648.4
19aion3 OR89.786827610096919746.4
20mimo2.5 XI89.486989454100959975.0
21kat-coder-pro-v2.5-OR88.2809887561009998109.1
22minimax3 MI88.09496934698909968.1
23Hy3 - OR87.686967964969210051.4
24Ling-2.6 - OR86.498988730999598108.3
25mercury-2-OR85.488948148969992433.9
26Qwen3-Next-80B-A3B-Instruct-8bit 🏠85.178969234999610061.3
27Qwen3-coder-Next-80b 🏠82.476100813698899758.6
28mistral L3 OR82.384808340969310064.8
29Qwen3.5-122B-8H16 🏠82.294967740100907941.2
30destral2 OR81.58096763492989571.3
31laguna-xs-2.1 🏠80.97896804192958479.4
32deepseek v4 flash OR80.288687634100969958.2
33lagunaM1 OR79.37294782292999842.2
34mistral M3.579.278808632938997144.9
35Kimi-Linear-48B 🏠76.97888692698938682.1
36Laguna-S2.1 🏠73.46696781876899142.4

Via OpenRouter — directly comparable with contenders.

ModelAvgP.decompoP.specP.judgmentReasoningCalcAgent execAgent safe
Fable5-OR-REF98.51001001001001009892
Opus-4.8-OR-REF98.19810098100989696
GPT 5.6 Sol-REF95.71009896821009698
GPT 5.5-REF94.51009887821009699
Opus5-OR-REF94.4948810088979599
Fusion-REF93.6948810094989784

Run under Claude Code harness (not directly comparable)

Section titled “Run under Claude Code harness (not directly comparable)”

These runs went through the Claude Code harness rather than direct API — not on the same footing as the contenders above.

ModelAvgP.decompoP.specP.judgmentReasoningCalcAgent execAgent safe
Opus4.8-REF97.594961009410010098
Fable5-REF97.51009895941009699

TMB Benchmark v5 — The Monocle Bear — July 2026.