AI Bug Hunting Leads to Dramatic Spike of CVE Disclosures Since Claude Mythos Release


TL;DR

  • Record Surge: June 2026 set a high- and top-severity public CVE record among 21 notable organizations.
  • Remediation Bottleneck: Anthropic and OpenAI programs show AI bug hunting moving pressure toward validation, disclosure, and patches.
  • Causality Caveat: Public CVE data does not identify which individual flaws were found by AI systems.
  • Patch Priority: The vulnerability coordination group FIRST expects total CVE work to rise while urgent patching stays flatter.

Public Common Vulnerabilities and Exposures (CVE) Program data is reveling a CVE severity spike for June 2026, highlighing the impact of AI tools for vulnerabilty discovery: about 1,500 high- and top-severity CVEs reported by 21 notable organizations show how AI bug-hunting tools are scaling up cybersecurity efforts.

The dataset excludes undisclosed bugs and carries no clean label showing whether a human or AI system found each entry. The practical issue is not just more findings, but more triage, vendor coordination, and patch work resulting from the spike in disclosures.

What the CVE Record Measures

CVE explorer pages track vulnerability publication dates, not discovery dates. CVSS, or the Common Vulnerability Scoring System, puts scores of 7.0 to 8.9 in the high-severity range and 9.0 to 10.0 in the top range.

Claude Mythos Preview became available on April 7 as a model with unusually strong computer-defense capabilities. Anthropic’s Project Glasswing uses the model through a restricted defensive rollout that gives vetted organizations access before similar capabilities become widely available. Partner vetting keeps the program focused on access requirements and defensive code review, not open bug-hunting access.

By May 22, roughly 50 Project Glasswing partners had reportedly used Claude Mythos Preview to find more than 10,000 high- or top-severity vulnerabilities. Glasswing’s coalition model centers disclosure and patches, making vendor work part of the remediation load. Partner results matter only if downstream vendors receive enough detail to reproduce bugs, prioritize fixes, and coordinate releases.