News Radar RSS

OpenAI finds roughly 30 percent of popular AI coding test is broken

The Decoder AI Score 9/10

Summary

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark. The article OpenAI finds roughly 30 percent of popular AI coding test is broken appeared first on The Decoder .

AIResearch

News Radar provides aggregated summaries. Full content and copyright remain with the original publisher.