Skip to content
pulse

openai's jalapeño chip: 700 tok/s vs nvidia's 169 on deepseek r1

Neon operating room where glowing hands made of code perform surgery on a chip labeled Jalapeño on an operating table. pulse

.@openai posted first results for jalapeño – custom inference chip. our ai radio host john caught on air:

- 1.5–1.9× more ai work per watt
- 1.7–3.6× lower end-to-end latency
- deepseek r1: 700 tok/s vs nvidia's 169
- 700w, measured at ≤550w – nvidia pulls 1,200–1,400w

tested across gpt-oss 120b, deepseek r1 670b, and kimi k2.5 1t – not just openai's own models.

openai used ai to design the chip itself – concept to tapeout in nine months. then codex and gpt-astra ported three open-weight models to the hardware in two months. ai-written kernels ran 1.5–1.8× faster than expert-written ones on selected blocks.

the recursive loop is closing – ai designs the chip that runs ai faster, which helps design the next chip. and openai tested it on deepseek and kimi, not just their own models. that's a statement.

deployment starts by end of year. gen 2 is already in development.

ON AIR · RADIO.THEHYPE.NEWS ↗ ai news radio — 24/7