Claude Opus 4.6, GPT-5.2 Score Only ~30% in New SWE-Atlas Benchmark
Article excerpt
Highlighted: the sentence this signal was extracted from
Claude Opus 4.6, GPT-5.2 score only ~30% in new SWE-Atlas benchmark. Scale AI's new benchmark evaluates AI coding agents across a spectrum of professional software engineering tasks. MARCH 4, 2026, 11:52 PM Scale AI has introduced SWE-Atlas, a new benchmark designed to evaluate how well AI coding agents perform real-world software engineering tasks inside complex codebases rather than simply generating code snippets. "SWE-Atlas is a benchmark for evaluating AI coding agents across a spectrum of professional software engineering tasks," the company said in its announcement. The benchmark includes three complementary leaderboards: Codebase QnA, Test Writing, and Refactoring. Of these, Codebase QnA is the first component released publicly, while the other two evaluations are expected to be introduced later. Codebase QnA focuses on testing how well AI agents understand large software systems before attempting modifications. The dataset contains 124 tasks drawn from 11 production repositories written in Go, Python, C, and TypeScript. Agents are placed inside sandboxed Docker environments containing the repositories and must answer technical questions by exploring the codebase, executing commands, and analysing runtime behaviour. Scale AI said that these tasks require running the software, tracing execution across multiple files, and synthesising findings. The benchmark...
Keep reading with a free account
The rest of this article, and every signal for Scale AI, is in your free account.
