At the start of September, Anthropic released two models at once: Claude Fable 5.1 and Claude Mythos 5.1. The naming is a little confusing, so here's the one sentence version from their announcement: they're the same model, with different safeguards.
- Fable 5.1 is what everyone gets. It's in the Claude apps, the API, and on AWS, Google Cloud, and Azure.
- Mythos 5.1 is the same model with looser safeguards for cybersecurity and life sciences work. You only get it through a vetted access program, and right now only some US organizations qualify.
If you read my post on what Claude Mythos means for your website, this is the next chapter. The capability that was locked away in April is now, in a limited form, available to anyone with an account.
What changed on the security side
Two things in the announcement matter here.
Fable 5.1 can now be used to find vulnerabilities. Anthropic's wording is that it can be used "to discover software vulnerabilities, though not to develop exploits for them." Before this, the safeguards were blunt enough that a lot of normal defensive work got blocked. Anthropic says the new safeguards produce 60% fewer false positives, and that people using Claude Code should see about 60% fewer interventions per session.
Mythos 5.1 is the strongest cyber model they've released. Anthropic says so directly, while also saying it still falls in the lower risk category of their own framework. It's available to vetted defenders through what they call the Cyber Verification Program. It also now powers Claude Security, their product that scans codebases for vulnerabilities and suggests patches.
The line between finding and exploiting
The rule for Fable 5.1 is basically: find bugs, don't weaponize them. That's a sensible line, and I'm glad they drew it. But think about what it means in practice.
Finding the bug is most of the work. Once you know there's an authentication bypass in a login endpoint, turning that into an attack usually isn't hard. And a model that won't write the exploit doesn't stop a person from reading the finding and writing it themselves, or from asking a different model that has fewer rules. Plenty of those exist.
So I'd read this as: guardrails on the big commercial models slow down casual misuse. They're not a wall. Assume that anyone who wants to find bugs in your app has good tools to do it.
The good news: you get the same tools
This cuts both ways, and the defender side is real.
If you have your code, you can now point a frontier model at it and ask it to look for security problems, and it'll actually do it instead of refusing. That's a big deal for small teams who've never had a security person. A few things I'd do:
- Ask it to review your auth and access control code specifically. Login, session handling, password reset, anything that checks roles or who owns what. That's where the scariest bugs in the Mythos write-up were.
- Have it look at every API route and ask what happens if a logged-in user calls it with someone else's ID. That one question finds a lot.
- Run it on dependency updates. When a library you use ships a security fix, ask whether your code uses the affected part.
What it can't do for you
A model reading your code sees what's in the repo. It doesn't see how the app is actually deployed: the headers your server sends, the admin panel someone left open on a subdomain, the .env file that's reachable over the web, the old API version still running behind the load balancer. Plenty of real problems live in the gap between the code and the running site.
That's what testing from the outside is for. Use the model on your code, then get someone independent to look at the live app the way an attacker would. The free header scan is a quick start. The full pen test is $495.