AI Models & LLM News: Releases, Benchmarks, Comparisons
New models, benchmarks and our tests
the performance gap between us and chinese models is at a record-low 6% per bloomberg
opus 5 re-enters openrouter top 10 at #8 – up 61% in a week
gemini 3.7 flash vs deepseek v4 pro 0813 vs muse spark 1.2 – on voxel city dioramas
glm 5.3 vs deepseek v4-pro: both saw the maze, neither escaped
nvidia is building a 1-trillion-parameter open model to beat deepseek and qwen
grok 4.6 vs gpt 5.6 sol: who builds better 3d castles in one html file
gpt 5.6 luna is the first us model in the openrouter top 3 since june. just in 2 weeks after a 80% price cut from outside the top 10
lfm2.5-vl-3b vs qwen3.5-4b: 3b beats 4.2b on visual grounding, but neither finds the button
muse glimmer 30b vs nemotron 3.5 lightning – on agentic coding
muse spark 1.2 vs gpt 5.6 sol vs kimi k3 vs grok 4.5 – on two landing pages