Anthropic’s fix for AI hacking: more AI
The company has introduced OSS Scanner, a free, opt-in service that runs periodic AI security scans on eligible open-source projects and sends maintainers detailed vulnerability reports.
With exploits now being built within minutes of a flaw surfacing, the company says projects that can find and fix bugs faster stand a better chance against attackers chasing the same weaknesses. Anthropic is putting its most powerful AI models, including Claude Mythos, to work on the open-source code that much of the internet is built on. The catch is that the reports are written entirely by AI. No human reviews or triages them before they reach a maintainer, and Anthropic admits some findings could be incorrect or invalid. Its argument is speed.
Anthropic is borrowing its criteria from Google’s OSS-Fuzz, which has scanned open-source software with fuzzers since 2016, so only projects with a critical impact on infrastructure and user security will qualify, decided case by case. Over the past six months, Anthropic’s models flagged more than 29,000 candidate vulnerabilities in major software, but its team could manually verify only about 6,000. Language models now catch over 85% of vulnerabilities on the CyberGym benchmark, up from under 20% at the start of 2025. Open-source projects are often run by unpaid volunteers, and the 2024 XZ Utils backdoor showed how a single compromised library could expose millions of systems.
Enrolment is not open to every repository. Core maintainers apply by submitting a pull request on GitHub. The service reportedly runs under Anthropic’s consumer terms, which allow inputs and outputs to be used for model training. OSS Scanner grows out of Project Glasswing. The stakes are familiar. Anthropic expects attackers to keep the upper hand for the next two years, until AI catches bugs before they ship.
Its penetration testers checked 97 critical and high-severity findings across 48 projects, and 85 of them, or 88%, met the bar for its coordinated disclosure process. Of the other 12, 11 were real bugs already reported elsewhere, leaving a single false positive. Each report comes with a self-contained reproducer, an explanation of the bug, a bisection showing when it entered the code where possible, and a candidate patch when one is available. Anthropic spent several weeks testing the pipeline with dozens of projects.

