Benchmark releases show capability gains but safety evals still lag
New benchmark results show continued gains in reasoning and coding, while independent safety evaluations struggle to keep pace with the speed of model releases.
2 items tagged AI Safety & Risks
New benchmark results show continued gains in reasoning and coding, while independent safety evaluations struggle to keep pace with the speed of model releases.
A structured comparison of frontier and open-weight language models through the lens of humanitarian and development deployment — capability, access, safety, cost and governance.