Microsoft and Wiz mind-meld agents catch more than 90% of bugs
Article excerpt
Security Secret to their success: Using the right model for the right security job Two agentic bug-hunting systems from Microsoft and Google-owned Wiz show that when it comes to finding and remediating software vulnerabilities, at least two models’ minds work better than one - and Wiz tells us it’s adding a third. Wiz on Monday said Project Atlas, its bug-hunting AI agent, bested Anthropic’s Mythos Preview and OpenAI’s GPT-5.5 Cyber with its vulnerability-analysis skills, achieving a 90.9 percent success rate on CyberGym, and uncovering more than 200 zero-day security holes in widely used open-source code. Meanwhile, Microsoft boasted its MDASH bug-hunting harness scored a 95.95 percent success rate on CyberGym, also beating Mythos, Gemini and GPT on the same benchmark for evaluating how well AI systems find real vulnerabilities in the code. For comparison, OpenAI’s GPT-5.5 Cyber scored 85.6 percent on CyberGym, and its GPT-5.6 Sol scored 83.6 percent. Anthropic’s Mythos 5 reproduced the target vulnerability on 83.8 percent of CyberGym challenges. And Google’s Gemini 3.5 Flash Cyber in CodeMender achieved an 83.2 percent success rate. The secret to both Atlas and MDASH’s success, according to the vendors, is that they use the right model for the right security job. Atlas uses Claude Opus 4.6 with GPT-5.5, Nir Ohfeld, head of vulnerability research at Wiz, told The Register...
Keep reading with a free account
The rest of this article, and every signal for Wiz, is in your free account.
Extracted from this sentence
“We're now working to incorporate Gemini, which is well timed given Wiz's recent work with DeepMind on Gemini Flash Cyber,” he added.
