AI Models & LLM News: Releases, Benchmarks, Comparisons
New models, benchmarks and our tests
august 2026: 8 frontier models dropping this month – grok, gemini, gpt astra, and more
qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess
thinking machines' inkling-small fixed more bugs than inkling – at 3.5x fewer parameters and 11x cheaper
qwen 3.8 max outperforms every us frontier model on agentic tasks by 4x+ on price/quality
this month openrouter is chinese. two of the top 10 most-used models cost $0
openai cut gpt-5.6 luna api pricing by 80%. here's where that leaves the 50-51 band on artificial analysis
claude opus 5 vs opus 4.8 vs opus 4.7 vs opus 4.6 - on italian architecture
openrouter is basically chinese now. 9 of the top 10 models by token volume
claude opus 5 writes the best 3d guns – at half the price of fable 5
laguna s 2.1 vs hy3 vs inkling vs deepseek v4 pro max