Cryptocurrency Prices by Coinlib

Microsoft Says MDASH Beats Claude Mythos and GPT-5.6 Sol in Cybersecurity Take a look at – Decrypt

Briefly
Microsoft says MDASH scored 95.95% on CyberGym, topping GPT-5.5 Cyber, Mythos 5, GPT-5.6 Sol, and Gemini 3.5 Flash Cyber.
MAI-Cyber-1-Flash handles as much as 90% of the workload, whereas MDASH sends the toughest instances to GPT-5.4.
The scanner is in non-public preview by way of Microsoft Defender, the place groups can assessment findings and generate proposed fixes.
Microsoft has launched its first devoted cybersecurity mannequin named MAI-Cyber-1-Flash and plugged it into MDASH, a vulnerability-hunting system that it says beats Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol whereas costing 50% lower than Microsoft’s present finest MDASH configuration, in keeping with the corporate.The mixed setup scored 95.95% on CyberGym, in keeping with Microsoft. CyberGym is a benchmark that asks AI brokers to breed 1,507 identified vulnerabilities throughout 188 open-source tasks, then scores them by the share efficiently reproduced in a managed surroundings.That put MDASH forward of GPT-5.5 Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 83.2%. The result's self-reported by Microsoft and had not appeared on CyberGym’s public leaderboard at publication time, although the benchmark makes use of a public take a look at set and an outlined success metric.MAI-Cyber-1-Flash doesn't work alone. Microsoft says it handles as much as 90% of duties, whereas MDASH routes the toughest 10% to GPT-5.4. That issues as a result of tokens—the chunks of textual content an AI processes—value cash each time a mannequin reads code or produces a solution.Picture: MicrosoftThis is the primary time a Microsoft mannequin constructed for effectivity is able to beating a dense cutting-edge mannequin constructed with basic capabilities in thoughts. “When mixed with MDASH, (MAI-Cyber-1-Flash) delivers world-class efficiency at 50 % of the price of main fashions,” Microsoft CEO Satya Nadella wrote.
As we speak, we're saying a sequence of updates that give prospects frontier-grade safety at half the fee.
MAI-Cyber-1-Flash is our first cybersecurity mannequin, constructed floor as much as discover probably the most difficult vulnerabilities in advanced code bases. When mixed with MDASH, it… pic.twitter.com/npcIihN1H7
— Satya Nadella (@satyanadella) July 27, 2026A mannequin (on this case MAI-Cyber-1-Flash) is the AI that causes over the code. A harness (MDASH on this case) is the equipment round it: the brokers, instruments, checks, and workflow that resolve the place to look, problem suspected findings, take away duplicates, and show {that a} bug will be triggered.MDASH makes use of greater than 100 specialised brokers assigned to audit code, debate whether or not a discovering is real, and construct a proof of idea—a working demonstration that the flaw exists.Microsoft stated occasional scans and delayed patches have gotten out of date as AI makes bug discovery cheaper. The corporate argues that many years of safety knowledge give it a bonus, including, “Nobody can manufacture this historical past.”Ever for the reason that launch of Claude Mythos, cybersecurity consultants have been making an attempt to beat or match its capabilities. Researchers reproduced Mythos-style vulnerability searching with public fashions for underneath $30 per scan. Dawid Moczadło, one of many researchers concerned, stated “the moat is transferring from mannequin entry to validation.”The scarce half is changing into the system that proves findings with out burying builders underneath false alarms.Decrypt additionally reported that GPT-5.5 Cyber had just lately taken the general public CyberGym lead with an 85.6% rating, narrowly beating Mythos. Microsoft’s rating is about 10 factors above GPT-5.5 Cyber and seven.5 factors above MDASH’s earlier outcome, but it surely evaluates the complete system reasonably than MAI-Cyber-1-Flash by itself.Microsoft is placing MDASH into non-public preview by way of Microsoft Safety Publicity Administration within the Defender portal. Clients can scan Git repositories, see findings ranked from unlikely to confirmed, and use the Defender CLI to generate proposed code fixes for developer assessment.The preview at present limits repositories to roughly 256MB and permits one concurrent scan per tenant. Challenge Notion is anticipated to increase the identical multi-agent method past code scanning into broader menace monitoring and remediation workflows.Every day Debrief NewsletterStart day-after-day with the highest information tales proper now, plus authentic options, a podcast, movies and extra.