AI Models & LLM News: Releases, Benchmarks, Comparisons
New models, benchmarks and our tests
Kimi K2.6 ships model + agent: 4,000 tool calls per run
z ai (glm) bans for non-coding tasks performed on their model
openai’s secret gpt pro upgrade looks like 5.5 “spud”
thehype compiled a top-6 ai models for the local stack
opus 4.7 is worse than opus 4.6 at spotting fakes
Opus 4.7: 50% More Tokens, 5h Limits, Devs Are Mixed
Claude Opus 4.7 Hits 87.6% on Agentic Coding, Beats 4.6
opus 4 to 4.7: the quiet shift no one’s pointing at