Skip to content

claude opus 5 vs opus 4.8 vs opus 4.7 vs opus 4.6 - on italian architecture

pulse Four arched doors labeled 4.6, 4.7, 4.8, and 5 with the last revealing Italian landmarks and a glowing key on cobblestones

context: opus 5 shipped jul 24 as anthropic's opus-tier model, not part of the mythos class – it outperforms its own flagship fable 5 on frontier-bench (43.3% vs 33.7%), gdpval-aa v2 (1861 vs 1747 elo) and osworld 2.0 (70.6% vs 66.1%), at half fable's price ($5/$25 vs $10/$50 per m tokens)

an opus-tier model outperforming the flagship on multiple benchmarks raised the question of how opus itself progressed across its version history, so we tracked it. the release cadence:

• opus 4.6 – feb 2026
• opus 4.7 – apr 2026
• opus 4.8 – may 2026
• opus 5 – jul 2026

then we ran our own test on the family itself: two icons of italian architecture, built from zero in single-file @threejs – no downloaded meshes, no image textures, every stone, arch and shadow generated in code, with a reference photo of the real landmark attached to each prompt to guide the generation

1. the leaning tower of pisa, modeled with its true curved lean axis from the mid-construction correction, not a straight tilt.

2. the colosseum as a ruin – true offset-curve arena, arc-length bay spacing, the exposed hypogeum grid

tested on @aimlapi across opus 5, opus 4.8, opus 4.7 and opus 4.6

- cost
#1 opus 4.7 – $7.41
#2 opus 4.8 – $9.80
#3 opus 5 – $10.19
#4 opus 4.6 – $13.96

- tokens
#1 opus 4.7 – 295k
#2 opus 4.8 – 392k
#3 opus 5 – 407k
#4 opus 4.6 – 558k

- lines of code
#1 opus 4.7 – 3,430
#2 opus 5 – 3,416
#3 opus 4.6 – 1,933
#4 opus 4.8 – 1,527

- generation time
#1 opus 4.8 – 31m
#2 opus 4.7 – 35m
#3 opus 4.6 – 39m
#4 opus 5 – 1h 33m

by the numbers: opus 4.7 is cheapest, uses the fewest tokens and produces the most code. opus 4.8 finishes fastest and produces the least code of the four. opus 5 takes the longest of the family – close to the combined time of the other three – and uses the most tokens and the highest cost

on accuracy: opus 5 was the only model of the four whose output matched the reference photos and dimensional specs on both prompts. opus 4.8, 4.7 and 4.6 showed larger deviations from the reference in the same comparison

conclusion: the rate per token hasn't moved – $5/$25 since opus 4.5, unchanged through 4.6, 4.7, 4.8 and 5. what's climbing is the cost per task: across the three most recent versions, combined spend on this test rises with each one – 4.7 at $7.41, 4.8 at $9.80, 5 at $10.19 – as each newer version uses more tokens to get there.

opus 4.6 is the exception, and the priciest of the four despite being the oldest

on accuracy, though, the pattern is clean: opus 5 was the only model of the four that matched the reference set on both landmarks

0:00
/0:49

Stay in the loop

Get the latest AI news delivered to your inbox weekly

Thanks for subscribing!