- likes
- 163
- comments
- 7
Post
Introducing CompanyBench, Turing's new long horizon benchmark and training dataset for enterprise knowledge work. What happens when frontier AI agents are given the same work real employees completed inside a company? We found out. CompanyBench evaluates AI models using five years of a fintech company's real operational history, including hundreds of database tables, thousands of files, and millions of Slack messages, emails, and Jira tickets. Every task is based on actual historical work completed by employees, with all private data thoroughly PII scrubbed. Across 64 challenging enterprise tasks, with each task run 10 times, today's leading models still struggle. The results: - GPT-5.5: 44% task completion - Claude Opus 4.8: 36% task completion We consistently observed four common failure patterns: - Guessing instead of consulting documentation - Searching for documents that don't exist - Applying filters that silently remove valid data - Producing the correct answer but failing to complete the final step, like posting to Slack or filing a ticket Enterprise AI needs more than strong reasoning. It needs the judgment, discipline, and thoroughness required to complete real work from start to finish. Read the full blog to explore the benchmark, methodology, and findings: https://lnkd.in/gd7U4G_i Interested in the dataset? Request samples here...
Keep reading with a free account
The rest of this post, and every signal for Turing, is in your free account.