@rachelmetz, ai reporter at @technology, broke down the cost math:
the industry is quietly moving from cost per token to cost per task. not how cheap the model is per million tokens – but how much it costs to actually get the job done.
companies and individual users are becoming cost-conscious fast. if you're doing computationally heavy tasks with a frontier model, it gets expensive quickly. chinese models offer what metz calls a really good bang for the buck.
bloomberg tested this directly: they asked models to build a website for a fictional coffee shop. fable 5 was the most expensive. chinese alternatives built passable websites at a 75% discount to claude.
kimi k3 is performing at benchmark levels comparable to the priciest us options on math, science and coding – built on a much humbler budget and with chips generations behind nvidia's lineup. the us-china performance gap narrowed from 9% in may to 6% in june per bloomberg intelligence.
chinese models now account for 41.4% of generative model downloads on hugging face – 5 points ahead of us models.
the old question was: which model is smarter? the new one: which model gets the job done for less?
full segment via link. source: @technology
the performance gap between us and chinese models is at a record-low 6% per bloomberg@rachelmetz, ai reporter at @technology, broke down the cost math:
— thehype. (@thehypedotnews) August 21, 2026
the industry is quietly moving from cost per token to cost per task. not how cheap the model is per million tokens – but how… pic.twitter.com/OE8Cqae6QR
Nick Trenkler