deepseek v4 pro built a working three.js halloween room for 4 cents. mistral large 4 built the same room from the same prompt for 61 cents. that is 14x the bill and more than twice the time.
large 4 is mistral's new 1-trillion-parameter mixture-of-experts model with 49 billion parameters active per token. it is in public preview, and mistral promises open weights for the end of october. we tested it on a real build job against three open models of the same size – deepseek v4 pro, kimi k2.6 and mimo v2.5 pro – and all four rooms render.
the models
| model | lab | total params | active params | weights |
|---|---|---|---|---|
| mistral large 4 | mistral | 1t | 49b | preview |
| deepseek v4 pro | deepseek | 1.6t | 49b | open |
| kimi k2.6 | moonshot ai | 1t | 32b | open |
| mimo v2.5 pro | xiaomi | 1.02t | 42b | open, mit license |
all four ran through openrouter.
the task
a halloween room with four objects:
- a snow globe
- a witch's cauldron
- a newton's cradle of skulls
- a crystal ball
the camera flies from one object to the next. same prompt for all four models.
the setup
the model plans first – room layout, camera path, shot timing and a file manifest with no file over ~350 lines – then writes the project one file per request, with everything it already wrote in context. our own harness on openrouter, 100k tokens max per reply, reasoning effort high for all four, a file that doesn't fit gets continued from its last line.
every room is a folder of plain es modules, three.js 0.170 from a cdn, every mesh and texture generated in code – no models, no images, no hdrs.
results
cost
| # | model | cost |
|---|---|---|
| 1 | deepseek v4 pro | $0.04 |
| 2 | mimo v2.5 pro | $0.16 |
| 3 | kimi k2.6 | $0.38 |
| 4 | mistral large 4 | $0.61 |
time
| # | model | time |
|---|---|---|
| 1 | deepseek v4 pro | 30m 09s |
| 2 | kimi k2.6 | 46m 43s |
| 3 | mistral large 4 | 72m 36s |
| 4 | mimo v2.5 pro | 72m 44s |
output tokens
| # | model | tokens |
|---|---|---|
| 1 | deepseek v4 pro | 83,539 |
| 2 | kimi k2.6 | 132,647 |
| 3 | mistral large 4 | 207,667 |
| 4 | mimo v2.5 pro | 234,014 |
lines of code
| # | model | loc |
|---|---|---|
| 1 | mistral large 4 | 3,467 |
| 2 | mimo v2.5 pro | 2,239 |
| 3 | deepseek v4 pro | 2,046 |
| 4 | kimi k2.6 | 1,868 |
observations
mistral large 4 wrote the most code, 3,467 lines across 14 files, and built the most detailed objects: a glass snow globe on a walnut base, a cast-iron cauldron on three legs with jars and a spellbook, an arched window with cobwebs in both corners. it is also the priciest: $0.61 for the room, and 72 minutes.
deepseek v4 pro is the cheapest and the fastest: 4 cents and 30 minutes for the whole room, with 83,539 output tokens, the fewest of the four.
kimi k2.6 has the cleanest wide shot of the whole table: the four objects, the window and the pumpkin piles on the floor all fit in one frame.
mimo v2.5 pro has the most atmosphere and the longest thinking: 72 minutes and 234,014 output tokens. warm candlelight everywhere, a row of candles along the table and a black cat on the windowsill.
our take
mistral large 4 builds the best-looking objects: its snow globe, cauldron and skull cradle are the ones you would put on a thumbnail. kimi k2.6 builds the room you can read at a glance, with the whole table in one frame. mimo v2.5 pro goes for mood over detail and lets the candlelight do most of the work. deepseek v4 pro's room is plainer and darker, but it is complete, and it cost 4 cents. mistral's extra detail comes at 14x that price.