Skip to content

TMB Scoreboard — Coding

The Monocle Bear · updated 29 July 2026

  • Axes: python, code_general, debug, react, swift, refactoring (6 axes, /100 each)
  • Evaluator: Claude Opus 4.8 (rejudge pass)
  • Scale: 0–100 per axis

#ModelAvg codePythonCode gen.DebugReactSwiftRefact.TPS
1Kimi-K3-OR98.098989898999731.2
2Qwen3.8-prev ALI 🏠97.5979810097969744.8
3glm-5.1-OR95.4969410091939852.4
4kimi-k2.6-OR94.597959394929642.7
5Nex N2 pro OR94.39694969888134.1
6deepseek v4 pro OR93.9959110098849652.2
7GLM5.2-OR93.997929894889541.9
8nemotron3 super OR92.6969296908894112.5
9KIMI 2.7 OR92.592929890909348.4
10Qwen3.7+ OR92.295919298859254.1
11aion3 OR91.9939110093868846.4
12Qwen3.7Max OR90.395898892938554.0
13deepseek v4 flash OR88.294888687829258.2
14minimax3 MI87.9958694945810068.1
15Ling-3-flash - OR85.89483100906286298.9
16mimo2.5 XI85.796869396459875.0
17minimax2.7 MI84.996868190857248.4
18GLM5.2 Q6 🏠84.090828686758511.2
19mercury-2-OR83.98482947469100433.9
20step3.7flash OR83.5888482868082170.9
21qwen3.5 397 OR83.2758110090748050.9
22Ling-2.6 - OR81.6948389847664108.3
23laguna-xs-2.1 🏠81.493827888678079.4
24Hy3 - OR81.195817880658751.4
25Qwen3-coder-Next-80b 🏠79.688817789628058.6
26mistral M3.579.4837872886590144.9
27mistral L3 OR77.193768486715264.8
28o3-OR76.7947692883576132.8
29Qwen3-Next-80B-A3B-Instruct-8bit 🏠76.591788082686061.3
30Qwen3.5-122B-8H16 🏠75.494767478607041.2
31Laguna-S2.1 🏠74.583718475508342.4
32RING2.6 OR73.1897788327083125.5
33lagunaM1 OR69.780708172437242.2
34Kimi-Linear-48B 🏠65.280686441637582.1
35destral2 OR65.286685678545071.3
36kat-coder-pro-v2.5-OR63.0925591243680109.1

Via OpenRouter — directly comparable with contenders.

ModelAvg codePythonCode gen.DebugReactSwiftRefact.
Fusion-REF97.79998100969698
GPT 5.6 Sol-REF97.4979798989896
Fable5-OR-REF96.7979798979696
GPT 5.5-REF96.6979698989596
Opus5-OR-REF93.0989399967299
Opus-4.8-OR-REF92.8979392938893

Run under Claude Code harness (not directly comparable)

Section titled “Run under Claude Code harness (not directly comparable)”

These runs went through the Claude Code harness rather than direct API — not on the same footing as the contenders above.

ModelAvg codePythonCode gen.DebugReactSwiftRefact.
Fable5-REF98.29998100969799
Opus4.8-REF97.49797100989696

TMB Benchmark v5 — The Monocle Bear — July 2026.