Skip to content
analysis

hy4 preview vs glm 5.3 vs qwen 3.8 max

Man holds a $3.38 receipt near wireframe skyscraper prints and a code-filled monitor, illustrating an ai coding benchmark. analysis

hy4 shipped all three for $0.537 – 3.2x under qwen. glm broke the fewest times, 2 fixes against hy4's 5. qwen wrote the most code, burned the most tokens, and still ships one of its three builds broken

the setup: three prompts of ~600 lines each, one html file per build, everything procedural – no meshes, no image files, no libraries past @threejs

prompts:

1. the petronas twin towers in dawn mist
2. taipei 101 in a tropical rainstorm
3. the cn tower in heavy snow

each file carries the geometry, the weather, night lighting, five scripted camera shots, a capture mode and a self-check panel that prints its own numbers. run through @OpenRouter, one shot per model, no reference images – two of the three models have no vision at all

models: @TencentHunyuan hy4 preview, @Alibaba_Qwen qwen 3.8 max, @Zai_org glm 5.3

- total cost, three builds
#1 hy4 preview – $0.537
#2 glm 5.3 – $1.111
#3 qwen 3.8 max – $1.737

- wall clock, three builds
#1 glm 5.3 – 44m
#2 hy4 preview – 50m
#3 qwen 3.8 max – 102m

- output tokens
#1 hy4 preview – 186,635
#2 glm 5.3 – 231,481
#3 qwen 3.8 max – 248,655

- fixes needed to make it run
#1 glm 5.3 – 2
#2 qwen 3.8 max – 4
#3 hy4 preview – 5

- lines of code shipped
#1 hy4 preview – 1,798
#2 glm 5.3 – 2,766
#3 qwen 3.8 max – 3,066

observations:

• every bug was one or two lines. no model failed the architecture – the geometry, the camera rigs and the self-check math were right everywhere. they broke on things a single run catches

• glm's petronas attempt spent 148,574 output tokens on reasoning and emitted zero characters of code. capping its thinking budget at 26k re-ran the same prompt in 683s for $0.250 – 3.2x faster and 2.7x cheaper

• no model won two towers in a row. taipei went to glm, petronas to qwen, cn tower back to glm, and the failures move the same way. the spread between tasks is bigger than the spread between models

conclusion: nine towers, 7,630 lines and 666,771 output tokens for $3.38 all in – and the cheapest model got there on 41% less code than the priciest!

watch the full video test via link

ON AIR · RADIO.THEHYPE.NEWS ↗ ai news radio — 24/7