Skip to content
analysis

gpt-6 sol vs grok 4.7 vs gpt-6 astra vs muse spark 1.3 – three star wars worlds each, one html file per world

Four hands cross purple, red, blue and orange lightsabers over a screen showing the Death Star, marking a four-model test analysis

the setup: one prompt, one reply, no agent loop. our own harness on @OpenRouter, headless chrome as the only judge – it loads the file, presses 1 / 2 / 3 / 4 and hands back a screenshot of every shot plus every console error. no rubric, no note from us. a file that crashed got one more turn with the console pasted back. grok 4.7 got extra turns on top, by our request – art direction as numbers, never a line edited by hand

every world is one html file, three.js 0.170 from a cdn, every texture generated in code – no model files, no images. an auto-playing cinematic with a shot timeline, letterbox and fades. reasoning high for astra, muse and sol; grok 4.7 ran at reasoning low

tasks:
1. death star – a slow orbital approach with star destroyers for scale, the corridor into the throne room, the station over a planet's horizon, then the superlaser: eight tributary beams, one green shot, a shockwave ring and thousands of instanced fragments
2. coruscant – the planet from orbit as a circuit board of glowing hubs, a daytime flythrough between towers, under a bridge and past the senate dome, a neon night district with a highway of light trails
3. kamino – an ocean world under storm clouds, tipoca city on stilts with gerstner waves, gpu rain and lightning, a glass walkway over a hall of marching clones

two hard rules in every brief: zero console errors on the first run, and nothing loaded from outside the file except three.js itself.

models:
@OpenAI gpt-6 sol, @xai grok 4.7, @OpenAI gpt-6 astra, @AIatMeta muse spark 1.3

all 12 worlds render. sol is the fastest on every one of the three tasks and never by less than 2x, and the whole grid came in at $6.99

total cost, three worlds
#1 muse spark 1.3 – $0.385
#2 gpt-6 sol – $0.556
#3 grok 4.7 – $1.732
#4 gpt-6 astra – $4.314

cost per world, sol against astra
death star – $0.202 vs $1.476
coruscant – $0.190 vs $1.559
kamino – $0.164 vs $1.279

wall clock, three worlds
#1 gpt-6 sol – 6m 23s
#2 muse spark 1.3 – 14m 27s
#3 gpt-6 astra – 26m 52s
#4 grok 4.7 – 58m 20s

total tokens
#1 gpt-6 sol – 59,299
#2 gpt-6 astra – 89,967
#3 muse spark 1.3 – 102,716
#4 grok 4.7 – 499,794

lines of code shipped
#1 gpt-6 astra – 4,274
#2 grok 4.7 – 4,034
#3 gpt-6 sol – 3,046
#4 muse spark 1.3 – 1,664

observations:

- gpt-6 sol shipped all three worlds on the first reply, 2m 21s or less each, zero console errors, 4k to 6k reasoning tokens per file. its coruscant is the only daytime city of the four with a bridge between the towers and traffic at three altitudes
- grok 4.7 is the model that takes direction. we gave it the exact camera path for tipoca city – 19 points with 2.5 units of clearance from every dome – and it landed the flythrough in one edit. we pointed at a one-character bug in its night window texture and it lit the whole district in one edit. 12 rounds across three worlds and not one console error in any of them
- gpt-6 astra is the closest to the film frames on the death star: a grey station with a dark side, the ring-walled throne room, a superlaser that cracks the planet before it blows. it is also 7.8x sol on cost
- muse spark 1.3 is the cheapest on every task and never by less than 1.2x against sol – $0.385 for the grid, 11x under astra
- sol did three worlds in 6m 23s, less than astra spent on any single one

watch the full test via link:

ON AIR · RADIO.THEHYPE.NEWS ↗ ai news radio — 24/7