we put the two models on one job: three cinematic three.js scenes, each a single self-contained html file, no textures, no models, no libraries beyond three.js. a rocket launch with stage separation, a meteor strike that knocks down a city, a dam break that washes away a town.
gpt-6 astra – @OpenAI, $5/$25 per m tokens on the flex tier via @OpenRouter
grok 4.7 – @SpaceXAI, shipped sep 21, $1.6/$4.8 per m tokens via @OpenRouter
- cost, three scenes
#1 grok 4.7 – $1.31
#2 gpt-6 astra – $4.11
- tries
#1 gpt-6 astra – 12
#2 grok 4.7 – 15
- wall clock, all calls summed
#1 gpt-6 astra – 25m 14s
#2 grok 4.7 – 60m 43s
- output tokens spent on thinking
#1 gpt-6 astra – ~34% (14–17k per scene)
#2 grok 4.7 – ~84% (62–86k per scene)
observations:
• grok thinks for 17–23 minutes per scene. 79k of its 93k rocket tokens were reasoning. a streaming request goes silent that long: one attempt was cut off and two more hung with zero bytes before a plain non-streaming call got through
• astra's first rocket had a concrete plain, we asked for grass, and a square bloom halo around the distant stage – both fixed by asking, not by editing. its meteor only switched the city lights off; "the buildings must physically collapse" was one more request
• grok's dam break rendered as a white ball. 6,200 spray particles drawn additively on top of a bloom pass, fixed by hand
grok 4.7 is 3x cheaper per scene and 2.4x slower
watch the full test via link
grok 4.7 vs gpt-6 astra in 3d scenes
— thehype. (@thehypedotnews) September 21, 2026
we put the two models on one job: three cinematic three.js scenes, each a single self-contained html file, no textures, no models, no libraries beyond three.js. a rocket launch with stage separation, a meteor strike that knocks down a city, a… https://t.co/9qOJbCyduN pic.twitter.com/SWrYrbQCrH