30 Mai

AI found 10000 critical bugs and the patches still are not keeping up

Anthropic published the first progress report on Project Glasswing on May 22. The numbers are staggering. In roughly 30 days, Claude Mythos Preview and about 50 partner organizations flagged 23,019 potential vulnerabilities in open-source software. Of those, 6,202 were classified as high or critical severity across more than 1,000 projects. Independent validation confirmed 1,726 as true positives. 1,094 earned a high or critical rating.

Only 97 have been patched.

That gap between discovery and remediation is the real headline. For decades the hard part of cybersecurity was finding bugs. Anthropic just proved that an AI model can surface them faster than the entire open-source ecosystem can process them. Some maintainers have asked Anthropic to slow down disclosures because they need more time to respond. Oracle shifted from quarterly to monthly patch releases. Microsoft warned its monthly patch volume will “continue trending larger for some time.”

The findings themselves are remarkable. A remote crash vulnerability in OpenBSD that had been hiding for 27 years. OpenBSD’s entire brand is security. Its tagline literally boasts about having only two remote holes in the default install. A 27-year bug in that codebase is not a minor miss. It proves that AI-assisted audit has fundamentally different reach than human review, even against the most paranoid codebases on earth.

Then there is the 16-year-old flaw in FFmpeg. FFmpeg processes video inside YouTube, Netflix, Zoom, Discord, VLC, Chrome, Firefox, and thousands of other applications. A flaw sitting in that code for 16 years means it was present across most of the streaming era.

And CVE-2026-5194 in WolfSSL, a CVSS 9.1 critical vulnerability in an embedded TLS library used in automotive systems, industrial controllers, and IoT devices. An attacker could forge certificates and impersonate legitimate services. On automotive and industrial systems. Nation-state threat actors build capabilities around exactly this kind of flaw. Glasswing found it first.

Cloudflare, one of the Glasswing launch partners, used Mythos Preview to scan its own infrastructure and found 2,000 bugs. 400 were high or critical. Cloudflare’s core business is internet security. Even they had a massive hidden vulnerability surface.

Mozilla fixed 271 vulnerabilities in Firefox through its Glasswing participation. Mozilla described this as a tenfold increase compared to findings from a prior Claude model. A 10x jump between model generations is not incremental improvement. It is a capability discontinuity.

IBM joined the consortium on May 19, bringing IBM Concert’s vulnerability management platform into the pipeline. The consortium now spans roughly 50 organizations including AWS, Apple, Google, Microsoft, Cisco, JPMorgan Chase, and Palo Alto Networks.

Anthropic is maintaining a strict 90-day coordinated vulnerability disclosure policy. Details of specific findings stay private during remediation. But the math is brutal. Anthropic has disclosed 530 high or critical bugs to maintainers so far. Only 75 are patched. Average fix time is about two weeks per bug. At that rate, the backlog grows faster than it shrinks.

The security industry has not built the tooling, automation, or human capacity to absorb vulnerability reports at AI speed. Finding bugs was supposed to be the hard part. It turns out that was the easy part.

Anthropic explicitly said no company, including itself, has developed safeguards strong enough to prevent models like Mythos from being misused. Mythos Preview remains restricted to vetted partners. OpenAI has a parallel program called Daybreak providing similar access to GPT-5.5-Cyber. Neither model is available to the public. Anthropic warned that models this capable will soon be developed by many AI companies, which makes the access control problem urgent in a way it was not six months ago.

Glasswing also demonstrated defensive use cases beyond bug hunting. One partner bank used Claude Mythos to detect and block a fraudulent $1.5 million wire transfer after an attacker breached a customer email account and made spoof phone calls. The model identified the fraud pattern before the transfer executed.

XBOW, an autonomous offensive security platform, described Mythos Preview as “a major advance” that is “substantially better than prior models at finding vulnerability candidates” and “adept at analysing source code with a security mindset.” Cloudflare noted the model excels at turning individual vulnerabilities into end-to-end attack chains, a capability that is equally valuable for defenders building threat models and dangerous in the wrong hands.

The bottleneck has shifted. The next phase of AI security is not about whether models can find bugs. They can. The question is whether the software ecosystem can fix them fast enough to matter.

Mots-cles

project glasswing claude mythos vulnerabilities anthropic cybersecurity patching bottleneck