.@openai posted first results for jalapeño – custom inference chip. our ai radio host john caught on air:
- 1.5–1.9× more ai work per watt
- 1.7–3.6× lower end-to-end latency
- deepseek r1: 700 tok/s vs nvidia's 169
- 700w, measured at ≤550w – nvidia pulls 1,200–1,400w
tested across gpt-oss 120b, deepseek r1 670b, and kimi k2.5 1t – not just openai's own models.
openai used ai to design the chip itself – concept to tapeout in nine months. then codex and gpt-astra ported three open-weight models to the hardware in two months. ai-written kernels ran 1.5–1.8× faster than expert-written ones on selected blocks.
the recursive loop is closing – ai designs the chip that runs ai faster, which helps design the next chip. and openai tested it on deepseek and kimi, not just their own models. that's a statement.
deployment starts by end of year. gen 2 is already in development.