The company that built its entire brand on AI safety just leaked its most dangerous model through an unsecured public data cache. Anthropic — the “responsible AI” startup founded by ex-OpenAI researchers who left because they thought OpenAI wasn’t careful enough — accidentally exposed draft blog posts, internal documents, and detailed descriptions of a secret new model called Claude Mythos. The model is, by Anthropic’s own admission, “far ahead of any other AI model in cyber capabilities” and poses “unprecedented cybersecurity risks.” The irony writes itself.

This isn’t a minor embarrassment. This is the AI safety company proving it can’t secure its own CMS, while simultaneously building a model it believes could break the internet’s defenses wide open. And the fallout has already started: cybersecurity stocks plunged, software stocks wobbled, and the entire AI industry is now grappling with a question nobody wanted to ask out loud — what happens when the models get too good at hacking?

A CMS Error Exposed 3,000 Unpublished Documents

Here’s how it happened. A “human error” in the configuration of Anthropic’s content management system left a draft blog post — and roughly 3,000 other unpublished assets — sitting in a publicly accessible, searchable data store. Fortune reviewed the documents first and broke the story on March 26. The draft post described a new model called Claude Mythos, part of a new tier called Capybara — larger and more intelligent than Anthropic’s existing Opus models, which were previously the company’s most powerful.

Think about that for a second. This is a company that routinely publishes 50-page safety reports about how carefully it tests its models before release. A company that makes “responsible scaling” its core marketing pitch. And it couldn’t lock down a data cache. It’s the equivalent of a bank vault manufacturer leaving the combination taped to the front door. Except the vault contains something far more dangerous than money.

Mythos Is Not Just an Upgrade — It’s a New Category

The leaked draft makes clear that Mythos isn’t an incremental improvement over Claude Opus 4.6. An Anthropic spokesperson confirmed the model represents “a step change” in AI performance and is “the most capable we’ve built to date.” The draft states that compared to Opus 4.6, the Capybara-tier model achieves “dramatically higher scores on tests of software coding, academic reasoning, and cybersecurity, among others.”

The naming convention matters here. Capybara is the tier — a new level above Opus in Anthropic’s model hierarchy. Mythos is the specific model name, like how “Opus 4” is a specific model within the Opus tier. This means Anthropic isn’t just releasing a better model; it’s creating an entirely new class of AI capability. The leaked draft described it as “by far the most powerful AI model we’ve ever developed.”

The model is currently in early access testing with select customers. It’s reportedly expensive to run and not ready for general availability — which, given what it can apparently do, might be the most responsible thing Anthropic has done in this entire saga.

The Cybersecurity Problem Is the Real Story

Forget the benchmarks for a moment. The most alarming part of the leaked draft is what Anthropic said about Mythos’s cyber capabilities. The company wrote — in its own words, in its own unpublished blog post — that the model is “currently far ahead of any other AI model in cyber capabilities” and “presages an upcoming wave of models that can exploit vulnerabilities in ways that far outpace the efforts of defenders.”

Read that again. Anthropic is telling us, inadvertently, that the offense-defense balance in cybersecurity is about to tip. Their own model can find and exploit software vulnerabilities faster than defenders can patch them. And they know it’s not just their model — it’s the beginning of a trend. Every frontier lab is building toward the same capabilities. Mythos just got there first.

The market reaction was immediate and brutal. CoinDesk reported that cybersecurity stocks plunged on the news, and software stocks followed. The logic is straightforward: if AI models can break software faster than humans can fix it, the entire cybersecurity industry’s value proposition changes overnight. Companies like CrowdStrike, Palo Alto Networks, and Fortinet sell the promise that they can keep you safe. What happens when the threat model includes an AI that can outrun every defender?

Follow the Money: Why Anthropic Built This Anyway

Here’s the part that nobody in the AI safety community wants to talk about. Anthropic knew this model posed unprecedented risks — the draft blog post says so explicitly. And they built it anyway. Not because they’re reckless, but because the economics of the AI race demand it.

Anthropic has raised over $15 billion in funding. Its investors — Google, Salesforce, and a parade of venture firms — didn’t write those checks for a company that stops at “good enough.” They wrote them for a company that builds the most powerful models in the world. The responsible scaling framework isn’t a brake; it’s a PR strategy that lets Anthropic push the frontier while maintaining the moral high ground over OpenAI and Google.

The Mythos leak reveals the tension at the heart of every “safety-first” AI company: you can’t be the safest lab and the most powerful lab at the same time. At some point, the capability curve outpaces the safety testing. And when your own draft blog post warns about “unprecedented cybersecurity risks,” you’ve crossed that line.

The Second-Order Effect: Every Government Just Got a Preview

This leak didn’t just embarrass Anthropic — it handed ammunition to every AI regulator on the planet. The EU’s AI Act is already in effect. China’s AI regulations are tightening. And in the U.S., the debate over AI regulation has been paralyzed by industry lobbying. But now there’s a document, in Anthropic’s own words, warning that a single AI model could “exploit vulnerabilities in ways that far outpace the efforts of defenders.”

Good luck lobbying against regulation when your own leaked blog post is the prosecution’s Exhibit A.

This also changes the calculus for national security agencies. If Anthropic’s model can do what the draft claims, then so can the next model from Google, OpenAI, or a Chinese lab with similar resources. The NSA, GCHQ, and their counterparts are now operating in a world where AI-powered cyber offense is about to become trivially accessible to any well-funded actor. The defense community has been warning about this for years. Now there’s proof — from the company that was supposed to be the careful one.

Who Gets Hurt: The Small Companies Without Defenses

When AI-powered exploitation goes mainstream, the victims won’t be Google or Microsoft — they have the resources to fight back. The victims will be mid-market companies, hospitals, municipal governments, and small businesses running outdated software with two-person IT departments. These organizations already struggle to keep up with human hackers. An AI that can scan for and exploit vulnerabilities at machine speed is a catastrophe waiting to happen.

The cybersecurity industry’s response will predictably be: “Buy more of our products.” But the Mythos leak suggests the gap between offense and defense is about to widen, not narrow. AI-powered defense tools exist, but they’re expensive, enterprise-focused, and years behind the offensive capabilities Anthropic just described. For the vast majority of organizations, the security posture just got significantly worse — and they don’t even know it yet.

The Verdict

Anthropic’s Mythos leak is the most consequential accidental disclosure in AI history. Not because the model is powerful — we all expected the next generation to be a step change. But because the company that positioned itself as the responsible adult in the room just proved two things simultaneously: it can’t secure its own infrastructure, and it’s building tools it considers a genuine threat to global cybersecurity.

The AI safety company that warns about existential risk from AI just created a model it describes as an unprecedented cyber weapon — and then left the description in an unlocked filing cabinet. If this doesn’t force a serious conversation about what happens when models get too capable, nothing will.

Mythos isn’t just Anthropic’s problem. It’s a preview of what every frontier lab is about to ship. And right now, nobody — not the regulators, not the defenders, not even Anthropic — has a plan for what comes next.