Skip to content

qwen 3.8 max outperforms every us frontier model on agentic tasks by 4x+ on price/quality

pulse 	Illustrated woman in patched leather jacket circles Qwen 35.0 in green, crosses out Opus and GPT scores on a whiteboard.

terminal-bench 2.1, blended api price per 1M tokens at a 3:1 input:output mix – qwen sits at $2.48 and scores 86.6

points of agentic capability per dollar:

• qwen 3.8 max: 35.0
• kimi k3: 14.2 – 2.5x behind
• claude opus 5: 8.9 – 3.9x behind
• gpt-5.6 sol: 8.0 – 4.4x behind
• claude fable 5: 4.2 – 8.3x behind

sol and opus 5 still score higher in absolute terms – 89.5 and 89.1. that's ~3 points, or 2-3 tasks (tb 2.1 is exactly 89 tasks, so points map almost 1:1 to tasks solved). you pay 4x for those

against fable 5 there's no tradeoff at all. qwen scores 86.6 to fable's 84.6 and costs 8x less. @Alibaba_Qwen said 3.8 max trails only fable 5 – on this benchmark it's ahead of it

the only model at the same price point is glm-5.2 at $2.15 blended. it scores 77.9. same money, 8.7 points apart

caveat: 86.6 is alibaba's own run, no independent number yet. everything else here is @ArtificialAnlys – terminus 2, e2b sandbox, pass@1 over 3 repeats. kimi self-reported 88.3 and independently came in at 85.0, so vendor numbers in this class have been running ~3 points hot

	Terminal-Bench 2.1 leaderboard: Qwen 3.8 Max hits 35.0 score per dollar, beating Kimi K3, Claude Opus 5, GPT-5.6 Sol and Claude Fable 5.

Stay in the loop

Get the latest AI news delivered to your inbox weekly

Thanks for subscribing!