Skip to content
opinions

nvidia says the best way to make tokens cheaper is increasing performance per watt. the market says: just use a cheaper model

Comic-style vendor with an NVIDIA badge shows a glowing gold watch beside cheap watches tagged $0.14 opinions

ian buck, nvidia's vp of hyperscale, at the ai infra summit:

agentic workloads are 100x heavier than 2023 chat – 142k input tokens, million-token kv caches, 65 agent turns. vera rubin delivers 30x throughput per megawatt vs blackwell on semianalysis agentx. the maxlps software fits 40% more gpus into the same power budget. lambda validated it: 19 nodes in the budget of 16.

but here's the tension nvidia doesn't address: while they optimize hardware to cut cost per token generation over generation, the token market itself is collapsing faster.

cost per million input tokens right now:

- deepseek v4 flash – $0.14
- gpt-5.6 luna – $0.20
- gpt-5.6 sol – $4
- claude opus 5 – $5
- claude fable 5 – $10

similar capability on coding benchmarks. 70x price spread.

nvidia's own customers are building around them:

- aws trainium 3 – claims 50% lower inference cost vs nvidia
- google tpu ironwood – inference-optimized, captive to gcp
- amd mi455x – 50% more memory per gpu, ships this year

every hyperscaler is designing custom silicon to reduce nvidia dependency.

nvidia's moat is the full stack – cuda, nvlink, annual cadence from vera rubin to feynman ultra. the question is whether the cost-per-token floor drops faster than nvidia can improve performance-per-watt. right now, the floor is winning.

ON AIR · RADIO.THEHYPE.NEWS ↗ ai news radio — 24/7