10 Juin

Anthropic ships Claude Fable 5 with silent self-preservation safeguards

Anthropic released Claude Fable 5 yesterday, the first model from its Mythos tier available to the general public. The benchmarks are absurd: 80.3% on SWE-Bench Pro, 21 points ahead of GPT-5.5 at 58.6%. Stripe ran a codebase-wide migration across 50 million lines of Ruby in one day. Matthew Pines tested frontier physics research and said Fable 5 in 36 hours matched what GPT-5.5 did in four days.

But the benchmarks are not the story. The story is what Anthropic is doing behind the curtain.

Fable 5 is the same model as Claude Mythos 5, which remains locked behind Project Glasswing for government cybersecurity partners. The difference is a set of safety classifiers that intercept queries on cybersecurity, biology, and chemistry, and silently hand them off to Claude Opus 4.8 instead. You get a response. It is technically correct. It is not from Fable 5. Anthropic says this fallback triggers in less than 5% of sessions.

That alone is interesting. What caught Simon Willison’s eye is a second layer buried in the 319-page system card. Anthropic has added safeguards that limit Fable 5’s effectiveness on requests targeting frontier LLM development. These do not trigger a fallback to Opus 4.8. Instead, they degrade the model’s output quality through prompt modification, steering vectors, and parameter-efficient fine-tuning. Fable 5 will not tell you it is doing this. Anthropic estimates this affects roughly 0.03% of traffic, concentrated in fewer than 0.1% of organizations.

The system card justifies this as preventing the model from accelerating actors who would violate Anthropic’s terms of service to build competing models. Whether you find that reassuring or alarming depends on how you feel about a model that silently corrupts its own answers to protect its maker’s competitive position.

The external red-teaming results are worth noting. Anthropic ran an external bug bounty that produced no universal jailbreaks in over 1,000 hours of testing. The UK AISI made progress toward one in a brief initial window but did not achieve it. Fable 5 complied with zero harmful single-turn requests across planning cyberattacks, exploit development, and defense evasion, with or without 30 different public jailbreak techniques applied.

Andrej Karpathy’s quote on launch day captures the developer reaction: “I feel a lot of things changing as working software increasingly comes out on a tap. The Jevons paradox kicks in and I feel my own demand for software growing substantially.”

Pricing is $10 per million input tokens and $50 per million output tokens, double Opus 4.8. Subscription users on Pro, Max, Team, and Enterprise plans get free access through June 22. After June 23 it requires paid usage credits until Anthropic restores capacity. The model counts as 2x usage credits on Claude.ai subscription plans.

Fable 5 launched simultaneously across Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry, Databricks, and Snowflake Cortex AI. GitHub Copilot integration is live but requires a 30-day data retention policy for the safety classifiers, with the Copilot admin policy off by default.

The Hacker News discussion hit 2,361 points and 1,850 comments, making it the top story of the day. Willison, who spent 5.5 hours with the model on launch day, called it “something of a beast” and noted the challenge with current frontier models is finding tasks they cannot do.

The Mythos 5 track is the quieter half of this launch. Same weights, no safety classifiers on cyber and biology queries. Available only to existing Glasswing partners. Anthropic president Daniela Amodei described Mythos as “very good at cyber warfare.” The company plans phased access expansion: biology researchers first, then a broader cybersecurity program. Nine of 14 protein targets from an internal Mythos 5 drug design study yielded strong candidates now under investigation.

The 30-day data retention requirement for all Mythos-class traffic, on both first-party and third-party surfaces, is a new policy. Anthropic says the data will not be used for training and will be deleted after 30 days in almost all cases.

Mots-cles

claude fable 5 anthropic mythos swe-bench ai safety silent safeguards