context: opus 5 shipped jul 24 as anthropic's opus-tier model, not part of the mythos class – it outperforms its own flagship fable 5 on frontier-bench (43.3% vs 33.7%), gdpval-aa v2 (1861 vs 1747 elo) and osworld 2.0 (70.6% vs 66.1%), at half fable's price ($5/$25 vs $10/$50 per m tokens)
an opus-tier model outperforming the flagship on multiple benchmarks raised the question of how opus itself progressed across its version history, so we tracked it. the release cadence:
• opus 4.6 – feb 2026
• opus 4.7 – apr 2026
• opus 4.8 – may 2026
• opus 5 – jul 2026
then we ran our own test on the family itself: two icons of italian architecture, built from zero in single-file @threejs – no downloaded meshes, no image textures, every stone, arch and shadow generated in code, with a reference photo of the real landmark attached to each prompt to guide the generation
1. the leaning tower of pisa, modeled with its true curved lean axis from the mid-construction correction, not a straight tilt.
2. the colosseum as a ruin – true offset-curve arena, arc-length bay spacing, the exposed hypogeum grid
tested on @aimlapi across opus 5, opus 4.8, opus 4.7 and opus 4.6
- cost
#1 opus 4.7 – $7.41
#2 opus 4.8 – $9.80
#3 opus 5 – $10.19
#4 opus 4.6 – $13.96
- tokens
#1 opus 4.7 – 295k
#2 opus 4.8 – 392k
#3 opus 5 – 407k
#4 opus 4.6 – 558k
- lines of code
#1 opus 4.7 – 3,430
#2 opus 5 – 3,416
#3 opus 4.6 – 1,933
#4 opus 4.8 – 1,527
- generation time
#1 opus 4.8 – 31m
#2 opus 4.7 – 35m
#3 opus 4.6 – 39m
#4 opus 5 – 1h 33m
by the numbers: opus 4.7 is cheapest, uses the fewest tokens and produces the most code. opus 4.8 finishes fastest and produces the least code of the four. opus 5 takes the longest of the family – close to the combined time of the other three – and uses the most tokens and the highest cost
on accuracy: opus 5 was the only model of the four whose output matched the reference photos and dimensional specs on both prompts. opus 4.8, 4.7 and 4.6 showed larger deviations from the reference in the same comparison
conclusion: the rate per token hasn't moved – $5/$25 since opus 4.5, unchanged through 4.6, 4.7, 4.8 and 5. what's climbing is the cost per task: across the three most recent versions, combined spend on this test rises with each one – 4.7 at $7.41, 4.8 at $9.80, 5 at $10.19 – as each newer version uses more tokens to get there.
opus 4.6 is the exception, and the priciest of the four despite being the oldest
on accuracy, though, the pattern is clean: opus 5 was the only model of the four that matched the reference set on both landmarks
On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art: pic.twitter.com/Cl3fDxM0gP
— Claude (@claudeai) July 24, 2026
Nick Trenkler