Benchmark releases show capability gains but safety evals still lag
New benchmark results show continued gains in reasoning and coding, while independent safety evaluations struggle to keep pace with the speed of model releases.
2 items tagged AI Impact
Open-weight frontier models are closing the gap with closed labs on core benchmarks, lowering the cost of private and on-device deployment for humanitarian and development organisations — while governance questions stay open.
New benchmark results show continued gains in reasoning and coding, while independent safety evaluations struggle to keep pace with the speed of model releases.