Skip to main content
Scale AILaunch

Claude Opus 4.6, GPT-5.2 Score Only ~30% in New SWE-Atlas Benchmark

What happened

Scale AI, Inc. launches SWE-Atlas.

Source

Article excerpt

Highlighted: the sentence this signal was extracted from

Claude Opus 4.6, GPT-5.2 score only ~30% in new SWE-Atlas benchmark. Scale AI's new benchmark evaluates AI coding agents across a spectrum of professional software engineering tasks. MARCH 4, 2026, 11:52 PM Scale AI has introduced SWE-Atlas, a new benchmark designed to evaluate how well AI coding agents perform real-world software engineering tasks inside complex codebases rather than simply generating code snippets. "SWE-Atlas is a benchmark for evaluating AI coding agents across a spectrum of professional software engineering tasks," the company said in its announcement. The benchmark includes three complementary leaderboards: Codebase QnA, Test Writing, and Refactoring. Of these, Codebase QnA is the first component released publicly, while the other two evaluations are expected to be introduced later. Codebase QnA focuses on testing how well AI agents understand large software systems before attempting modifications. The dataset contains 124 tasks drawn from 11 production repositories written in Go, Python, C, and TypeScript. Agents are placed inside sandboxed Docker environments containing the repositories and must answer technical questions by exploring the codebase, executing commands, and analysing runtime behaviour. Scale AI said that these tasks require running the software, tracing execution across multiple files, and synthesising findings. The benchmark...

Keep reading with a free account

The rest of this article, and every signal for Scale AI, is in your free account.

Extracted by Autobound

From the Signal API record
Event
Launch

What this signalsA launch often needs new go-to-market and support spend.

Product
SWE-Atlas

The full record

From the Signal API record

Details

Category
Launches

Extraction

Confidence
90%
Detected
Mar 5, 2026
signal_type
news
signal_subtype
launches

Use this data

Get every launch signal for Scale AI and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Scale AI this week?”

  2. Send it to your own tools

    The Signal API returns launch signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full news record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/b93b6486-bf3f-426c-ad7e-9e3386e0fb0b returns this record as JSON. POST /v1/companies/enrich returns every signal for scale.com.

{
  "signal_id": "b93b6486-bf3f-426c-ad7e-9e3386e0fb0b",
  "signal_type": "news",
  "signal_subtype": "launches",
  "detected_at": "2026-03-05T07:53:41+00:00",
  "company": {
    "name": "Scale AI",
    "domain": "scale.com"
  },
  "data": {
    "url": "https://analyticsindiamag.com/ai-news/claude-opus-46-gpt-52-score-only-30-in-new-swe-atlas-benchmark",
    "title": "Claude Opus 4.6, GPT-5.2 Score Only ~30% in New SWE-Atlas Benchmark",
    "excerpt": "Claude Opus 4.6, GPT-5.2 score only ~30% in new SWE-Atlas benchmark.\n\nScale AI's new benchmark evaluates AI coding agents across a spectrum of professional software engineering tasks.\n\nMARCH 4, 2026, 11:52 PM\n\nScale AI has introduced SWE-Atlas, a new benchmark designed to evaluate how well AI coding agents perform real-world software engineering tasks inside complex codebases rather than simply generating code snippets.\n\n\"SWE-Atlas is a benchmark for evaluating AI coding agents across a spectrum of professional software engineering tasks,\" the company said in its announcement.\n\nThe benchmark includes three complementary leaderboards: Codebase QnA, Test Writing, and Refactoring.\n\nOf these, Codebase QnA is the first component released publicly, while the other two evaluations are expected to be introduced later.\n\nCodebase QnA focuses on testing how well AI agents understand large software systems before attempting modifications.\n\nThe dataset contains 124 tasks drawn from 11 production repositories written in Go, Python, C, and TypeScript. Agents are placed inside sandboxed Docker environments containing the repositories and must answer technical questions by exploring the codebase, executing commands, and analysing runtime behaviour.\n\nScale AI said that these tasks require running the software, tracing execution across multiple files, and synthesising findings.\n\nThe benchmark...",
    "product": "SWE-Atlas",
    "summary": "Scale AI, Inc. launches SWE-Atlas.",
    "category": "launches",
    "found_at": "2026-03-05T07:53:41Z",
    "planning": false,
    "image_url": "https://storage.googleapis.com/gpt-engineer-file-uploads/XW65pun0vjQL3xcUmVglEg4ZrsL2/social-images/social-1763560217027-Expert-Panel-discussion-768x512.webp",
    "confidence": 0.8973,
    "product_data": {
      "full_text": "SWE-Atlas",
      "fuzzy_match": true
    },
    "published_at": "2026-03-05T07:53:41Z",
    "article_sentence": "Scale AI has introduced SWE-Atlas, a new benchmark designed to evaluate how well AI coding agents perform real-world software engineering tasks inside complex codebases rather than simply generating code snippets."
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.