Skip to content

mai image 2.6 vs grok imagine image 2.0 vs gpt image 2 vs reve 2.1 – on text-to-image

pulse 	Illustrated robots sitting at school desks writing on paper under a chalkboard that reads TEXT

@microsoftai just shipped mai-image-2.6, second on the @arena text-to-image leaderboard, +79 elo over 2.5 overall and +91 elo on text rendering alone. they say it puts mai ahead of @aiatmeta, @google and @spacexai. for now arena is the only place you can use it – mai playground gets it later this week, microsoft foundry and the rest of their products after that

so we built a 5-prompt set aimed straight at what they claim:

1. a cutaway 3d render of a nordic house with both floors visible – 3d modeling

2. a soda can frozen mid-splash in a juice explosion – commercial design

3. a cereal box on a shelf with a full nutrition panel, ingredients, barcode and lot number in fine print – text rendering

4. a stop-motion frame with three characters felted from wool – cartoon

5. a boxer in her corner, sweat, blood and hands in frame – portraits

same prompts, one attempt each, no cherry-picking, four models

generation time per image:

• house – mai 15s / grok 20s / gpt 54s / reve 17s
• soda can – mai 41s / grok 22s / gpt 59s / reve 27s
• cereal box – mai 41s / grok 18s / gpt 57s / reve 34s
• wool frame – mai 38s / grok 23s / gpt 62s / reve 30s
• boxer – mai 30s / grok 18s / gpt 51s / reve 25s

observations:

• the +91 elo on text is real. the cereal side panel is roughly ninety characters of 6pt print and mai is the only model that renders all of it correctly – nutrition table, ingredients, phone number, barcode digits, best before date and lot number. reve collapsed into fake words and mirrored letters, grok garbled the ingredients, gpt held on but softened

• mai also wins on physical logic. the house asked for a stove flue rising through both floors – mai runs one continuous pipe from the stove to the roof, while grok and reve each drew two chimneys, an interior flue plus a decorative masonry stack connected to nothing

• in the wool scene mai is the only one where the knocked-over cup is actually spilled and the cat is reacting to it. everyone else placed the cup as a prop

• the flip side of that text training: mai adds copy nobody asked for. it turned the soda shot into a finished ad layout with an invented headline and a slogan, which is a miss – the brief was a photograph, not a magazine spread. cleanest execution there was grok, exact label text, no additions, best splash

• nobody knows what an enswell is. the boxer prompt named the tool by name and grok produced a flat metal plate with no handle, mai a handled piece shaped like a stamp, gpt and reve just a folded napkin

• grok is the speed story again – 18 to 23 seconds across the set and it never collapses. gpt image 2 is the slowest by a wide margin at 51 to 62s and took no prompts outright

overall: mai-image-2.6 takes four of five. the arena gains hold up where they're measured, and the 3d category they scored highest on is exactly where it beat everyone on structure rather than looks. the one loss is instruction following – trained so hard on text that it can't stay quiet when the brief doesn't ask for any

new favorite for anything with type in the frame. grok still owns speed

watch the full test via link

Stay in the loop

Get the latest AI news delivered to your inbox weekly

Thanks for subscribing!