Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, calling them the world’s most advanced models for coding and knowledge work. The announcement drew close to 2 million views within hours, one week after rumors about OpenAI’s unreleased Astra dominated the conversation.
The release is less about raw intelligence than about endurance and economics. Fable 5.1 more than doubles its predecessor on a scientific research benchmark, holds its price, and cuts the cost of the most common expensive workloads by a quarter to nearly half. Here is what changed, what it costs, who gets Mythos, and what it means if you are not an engineer.
What Anthropic Announced
Two names, one underlying model. Claude Fable 5.1 is the generally available version, running with Anthropic’s production safeguards, and it is live today across the Claude apps, Claude Code, Claude Enterprise, the API, and Amazon Bedrock, Google Cloud, and Microsoft Foundry. Claude Mythos 5.1 is the same model with more permissive safeguards for verified cybersecurity defenders and life scientists, available only through trusted access programs. Both are documented on Anthropic’s official announcement page.
Specifications at a glance
| Spec | Claude Fable 5.1 |
|---|---|
| Context window | 1 million tokens |
| Maximum output | 128,000 tokens |
| Inputs and output | Text and images in, text out |
| Thinking | Adaptive, always on, default effort high |
| Knowledge cutoff | June 2026 |
| API model ID | claude-fable-5-1 |
Mythos 5.1 shares the same specifications and pricing.
Claude Fable 5.1 vs Fable 5: The Benchmarks
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Agentic scientific research (Terminal-Bench-Science 0.1) | 52.6% | 24.7% | 29.0% | 22.4% |
| Agentic coding (Terminal-Bench 4.0) | 55.8% | 42.0% | 52.3% | 37.3% |
| Knowledge work (GDPval-AA v2) | 1853 | 1723 | 1824 | 1711 |
| Business workflows (AutomationBench) | 31.4% | 17.1% | 26.9% | 19.6% |
| Agentic coding (CursorBench 3.2.0) | 73.4% | 70.5% | 70.0% | 67.2% |
| Computer use (OSWorld 2.0, strict) | 41.7% | 36.1% | 39.6% | n/a |
| Multidisciplinary reasoning (Humanity’s Last Exam, with tools) | 65.0% | 63.8% | 63.6% | n/a |
Anthropic says Fable 5.1 was evaluated with its production safeguards enabled, and reports a standard error of 3.5 to 4.5 points on Terminal-Bench-Science, so small differences on that benchmark should be read cautiously.
Where the jump is biggest
The pattern is clear: the longer and more autonomous the task, the bigger the gain. Scientific research more than doubled, business workflow automation improved 84 percent in relative terms, and Terminal-Bench coding rose almost 14 points; Mythos 5.1 reaches 60.9 percent on the same benchmark.
Where it is modest
On short-horizon tests the picture is incremental: three points on CursorBench, about one point on Humanity’s Last Exam. One widely shared analyst summary put it fairly: dramatic gains on long-running tasks, modest improvements elsewhere, not a uniform intelligence jump.
Claude Fable 5.1 Pricing: Same Rates, Much Cheaper in Practice
| Price component | Fable 5.1 | Fable 5 |
|---|---|---|
| Input | $10 per million tokens | $10 |
| Output | $50 per million tokens | $50 |
| Cache reads | $0.25 per million tokens | $1.00 |
| Cache writes | $12.50 (5-minute) / $20 (1-hour) | same |
| Batch processing | $5 input / $25 output | same |
The one changed line matters more than it looks. Cache reads, the repeated re-reading of context that dominates agentic work, cost 75 percent less, which Anthropic says lowers Fable’s effective cost by about 25 percent on typical workloads and up to 45 percent on context-heavy agentic ones, measured over four weeks of August usage. One caveat early users flagged: the saving applies where usage is billed by the token, so subscribers do not see a 45 percent discount.
The effort-level trick
Fable 5.1 exposes five effort levels, and the cost curve is the most interesting chart in the release: at the lowest effort the model scores about 26 percent on the science benchmark at roughly $11 per task, while Fable 5 at maximum effort scored 24.7 percent at roughly $44. The cheapest setting of the new model beats the most expensive setting of the old one at a quarter of the price, and maximum effort climbs to 52.6 percent for around $38. Choosing the right effort level is now a real lever for anyone paying by the token.
What Changed in Everyday Use
Anthropic’s developer team summarized the practical differences: the model gets much further into a long task before needing your input, is better at saying when it is stuck, and writes more naturally. Early testers echo the efficiency: financial research firm Rogo reported equal accuracy with 20 percent fewer tokens, Red Hat found updates more concise, and media company Every said its Slack agent used under half as many tokens as Opus 5 at about twice the speed.
In one individual demonstration shared on X, an Anthropic staffer gave the model a photo of an empty lot; it designed a house, rendered it, and produced a cinematic video walkthrough through code.
Fable 5.1 and Mythos 5.1 as Research Tools
The scientific examples are the announcement’s boldest claims, and two of the three come from the restricted Mythos 5.1. In molecular design, Mythos 5.1 used open-source protein tools, with binders validated in two external laboratories: the hit rate reached close to 50 percent across 12 targets, against a typical 10 to 15 percent, and on three targets the binders outperformed the best designs from Adaptyv Bio’s protein design competitions. In planetary science, Fable 5.1 trained a neural network on radar images from NASA’s Magellan mission, taken more than 30 years ago, to map a third of Venus at 2 to 3 kilometer resolution instead of 10 to 20, with heights up to 25 percent more accurate. In computational biology, Mythos 5.1 wrote custom GPU kernels that sped up seven open-source deep learning models by up to 2.5 times, cutting estimated GPU costs by 30 to 60 percent on genome-wide analyses, and Anthropic says it plans to open-source those optimizations.
Enterprise stories point the same way. Millennium said the model traced a rare crash to a bug in an external vendor library that had resisted explanation for four to five years. Ramp documented a 38-hour unattended machine learning run in which the model re-evaluated results, launched new experiments, and returned findings. Browserbase reported 82 percent completion on its hardest browser-agent tasks, against 74 for Opus 5 and 57 for Fable 5. These are company-provided examples, not independent benchmarks.
What Is Claude Mythos 5.1?
Mythos 5.1 uses the same underlying model as Fable 5.1 but applies more permissive safeguards for vetted cybersecurity and life-sciences users. It is not an unrestricted model: Anthropic’s other safeguards and usage policy remain in force, and it is not a consumer product. Access runs through two invitation-only programs under what Anthropic’s documentation calls Project Glasswing: the Cyber Verification Program for defensive security professionals and the Life Sciences Verification Program, developed with the US government. Both are currently limited to US organizations, with international expansion coordinated with government partners. The 60.9 percent coding score shows what the model can do under those more permissive safeguards; restricted access is how Anthropic squares that with its safety commitments.
Fewer False Alarms: The Safeguard Changes
For ordinary users, the quieter change may matter most. Cybersecurity safeguards now flag benign requests about 60 percent less often per Claude Code session, and defensive vulnerability discovery is permitted, while exploit generation, penetration testing, and binary scanning remain restricted. On basic biology and medical questions, fallback to a refusal or weaker answer has dropped by around 85 percent; anyone who has hit that wall will notice.
Enterprise Frontier Safeguards Explained
For large companies, the most consequential part is Enterprise Frontier Safeguards. It solves a real dilemma: detecting sophisticated misuse requires retaining activity data to spot cross-session patterns, while regulated enterprises demand zero data retention. EFS can keep activity logs in the customer’s own cloud storage, optionally under customer-managed encryption keys, and run automated pattern analysis without Anthropic employees reviewing the data, with each control opt-in; flagged patterns go to the customer’s security team. There is no fee beyond the customer’s cloud costs, more than 100 enterprises helped design it, it rolls out in phases this fall, and eligible customers get zero data retention meanwhile.
How It Compares: Opus 5, GPT-5.6 Sol, and the Astra Question
Against Anthropic’s own Opus 5, Fable 5.1 leads on every benchmark in the launch table, and one coding company announced it was moving its Opus 5 traffic to the new model on launch day. Against OpenAI’s GPT-5.6 Sol, per-token pricing looks unfavorable, since Sol’s promotion runs $4 input and $20 output through late November, but the benchmark gaps are wide: 22.4 versus 52.6 percent on scientific research. Anthropic’s argument is task-completion economics: a model that finishes the job in fewer tokens and fewer retries can be cheaper per result while costing more per token.
The unspoken comparison is OpenAI’s unreleased Astra; shipping a broadly available model with published numbers while Astra remains rumor is a statement in itself.
What Early Users Are Saying
The praise clusters around endurance and honesty: it runs longer without hand-holding and admits when it is stuck. Trading firm Jane Street said it solves more problems than Fable 5 or Opus 5; MongoDB described a complex prototype built in three days unattended. The criticism is about who benefits: the pricing gains flow to API and enterprise customers, and one popular summary complained the release favors business users over subscribers. There is also healthy skepticism about reading a doubling on one benchmark, with a 4-point standard error, as a general leap.
What It Means If You Are Not an Engineer
Three things carry over. First, longer autonomy: the model gets further into a multi-step task before asking you anything, which makes delegation, not prompting, the skill that pays. Second, effort levels, choosing how hard the model should think for a job, are becoming a standard control, and knowing when low effort is enough is a money skill. Third, fewer needless interventions on health and science basics make the assistant more useful for the questions non-experts ask most.
The Skills That Outlast Every Model Version
Fable 5.1 is one release in a crowded season. The durable advantage belongs to people who can describe a task precisely, decide how much autonomy to grant, check the result, and switch tools without relearning everything. Coursiv builds those foundations with step-by-step guides, short daily lessons, and hands-on practice with AI tools, designed for busy people without a technical background, so a new model version becomes an upgrade rather than a restart. Check the official site for current course details and pricing.
What to Watch Next
Watch the EFS rollout this fall, a possible template for regulated industries. Watch whether the Mythos access programs expand beyond the US. And watch OpenAI’s response: if Astra ships comparable long-horizon numbers at lower token prices, effort levels and caching economics become the battleground.
The Bottom Line
Claude Fable 5.1 is an endurance upgrade more than an intelligence upgrade, and that is the more useful kind right now. It works longer, can cost less per finished task on the workloads Anthropic measured, triggers fewer needless safeguard interventions, and comes with a privacy architecture that enterprises actually asked for. The models keep changing; the ability to put them to work is what compounds.