Skip to main content
TuringLinkedIn

Turing introduces CompanyBench, a benchmark dataset for enterprise AI agents using real operational data.

Source

LinkedInAug 11, 2026
likes
163
comments
7

Post

Introducing CompanyBench, Turing's new long horizon benchmark and training dataset for enterprise knowledge work. What happens when frontier AI agents are given the same work real employees completed inside a company? We found out. CompanyBench evaluates AI models using five years of a fintech company's real operational history, including hundreds of database tables, thousands of files, and millions of Slack messages, emails, and Jira tickets. Every task is based on actual historical work completed by employees, with all private data thoroughly PII scrubbed. Across 64 challenging enterprise tasks, with each task run 10 times, today's leading models still struggle. The results: - GPT-5.5: 44% task completion - Claude Opus 4.8: 36% task completion We consistently observed four common failure patterns: - Guessing instead of consulting documentation - Searching for documents that don't exist - Applying filters that silently remove valid data - Producing the correct answer but failing to complete the final step, like posting to Slack or filing a ticket Enterprise AI needs more than strong reasoning. It needs the judgment, discipline, and thoroughness required to complete real work from start to finish. Read the full blog to explore the benchmark, methodology, and findings: https://lnkd.in/gd7U4G_i Interested in the dataset? Request samples here...

Keep reading with a free account

The rest of this post, and every signal for Turing, is in your free account.

Extracted by Autobound

From the Signal API record
Signal
LinkedIn

What this signalsCompany posts often show what the team is pushing right now.

The full record

From the Signal API record

Topics and mentions

Tags

  • Artificial Intelligence
  • Data Science
  • Product Development
  • Enterprise Software
  • Data Management

Initiatives

  • developing benchmark for enterprise AI agents
  • creating training dataset for enterprise knowledge work
  • evaluating leading AI models on enterprise tasks

Pain points

  • AI agents struggle with enterprise knowledge work tasks
  • AI models fail to complete real-world enterprise tasks
  • AI models exhibit common failure patterns like guessing and document searching

Technologies named

  • GPT-5.5
  • Claude Opus 4.8
  • Slack
  • Jira

Extraction

Detected
Aug 13, 2026
signal_type
linkedin-post-company
signal_subtype
linkedinPost

Use this data

Get every LinkedIn signal for Turing and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Turing this week?”

  2. Send it to your own tools

    The Signal API returns LinkedIn signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full linkedin-post-company record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/8d920be3-4e81-4492-9b07-faf1f739aad1 returns this record as JSON. POST /v1/companies/enrich returns every signal for turing.com.

{
  "signal_id": "8d920be3-4e81-4492-9b07-faf1f739aad1",
  "signal_type": "linkedin-post-company",
  "signal_subtype": "linkedinPost",
  "detected_at": "2026-08-13T14:27:29.626+00:00",
  "company": {
    "name": "Turing",
    "domain": "turing.com"
  },
  "data": {
    "tags": [
      "Artificial Intelligence",
      "Data Science",
      "Product Development",
      "Enterprise Software",
      "Data Management"
    ],
    "summary": "Turing introduces CompanyBench, a benchmark dataset for enterprise AI agents using real operational data.",
    "post_url": "https://www.linkedin.com/posts/turingcom_introducing-companybench-turings-new-long-activity-7490791593600311296-iRLn",
    "num_likes": 163,
    "post_text": "Introducing CompanyBench, Turing's new long horizon benchmark and training dataset for enterprise knowledge work.\n\nWhat happens when frontier AI agents are given the same work real employees completed inside a company?\n\nWe found out.\n\nCompanyBench evaluates AI models using five years of a fintech company's real operational history, including hundreds of database tables, thousands of files, and millions of Slack messages, emails, and Jira tickets. Every task is based on actual historical work completed by employees, with all private data thoroughly PII scrubbed.\n\nAcross 64 challenging enterprise tasks, with each task run 10 times, today's leading models still struggle.\n\nThe results:\n- GPT-5.5: 44% task completion\n- Claude Opus 4.8: 36% task completion\n\nWe consistently observed four common failure patterns:\n- Guessing instead of consulting documentation\n- Searching for documents that don't exist\n- Applying filters that silently remove valid data\n- Producing the correct answer but failing to complete the final step, like posting to Slack or filing a ticket\nEnterprise AI needs more than strong reasoning. It needs the judgment, discipline, and thoroughness required to complete real work from start to finish.\n\nRead the full blog to explore the benchmark, methodology, and findings: https://lnkd.in/gd7U4G_i \n\nInterested in the dataset? Request samples here...",
    "initiatives": [
      {
        "topic": "developing benchmark for enterprise AI agents",
        "urgency": 0.9
      },
      {
        "topic": "creating training dataset for enterprise knowledge work",
        "urgency": 0.9
      },
      {
        "topic": "evaluating leading AI models on enterprise tasks",
        "urgency": 0.8
      }
    ],
    "pain_points": [
      {
        "topic": "AI agents struggle with enterprise knowledge work tasks",
        "intensity": 0.8
      },
      {
        "topic": "AI models fail to complete real-world enterprise tasks",
        "intensity": 0.75
      },
      {
        "topic": "AI models exhibit common failure patterns like guessing and document searching",
        "intensity": 0.65
      }
    ],
    "posted_date": "2026-08-11T20:41:28.542Z",
    "num_comments": 7,
    "technologies_mentioned": [
      {
        "name": "GPT-5.5",
        "status": "evaluated"
      },
      {
        "name": "Claude Opus 4.8",
        "status": "evaluated"
      },
      {
        "name": "Slack",
        "status": "using"
      },
      {
        "name": "Jira",
        "status": "using"
      }
    ]
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.