Lotu Radar About

OpenAI finds roughly 30 percent of popular AI coding test is broken

The Decoder AI Score 9/10

Summary

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark. The article OpenAI finds roughly 30 percent of popular AI coding test is broken appeared first on The Decoder .

AIResearch

Lotu Radar provides attributed news summaries and links to the original publisher. Full reporting and copyright remain with the source.