Anthropic:Newsroom(网页) · 2026/9/2

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. Claude Fable 5.1 an

We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences. Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards. Price . Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%. Data retention . Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention. Safeguards . We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon. A new performance frontier Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.) Agentic scientific research Agentic terminal coding Multidisciplinary reasoning Agentic coding Agentic scientific research Agentic terminal coding Multidisciplinary reasoning Agentic coding Terminal-Bench-Science 0.1 Accuracy vs Cost Fable 5.1 Fable 5 Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise. Terminal-Bench 4.0 Accuracy vs Cost Mythos 5.1 Fable 5.1 Mythos 5 Terminal-Bench 4.0 scores by cost (log scale), at each effort level. Claude Fable 5.1 and Claude Mythos 5.1 are the same underlying model; the gap between them reflects the tasks on which our earlier, less precise cyber safeguards intervened. With the improvements we’re making to these safeguards today, we expect the difference between the models to be much smaller. Humanity's Last Exam Accuracy vs Cost Fable 5.1 (with tools) Fable 5.1 (no tools) Fable 5 (with tools) Fable 5 (no tools) Humanity’s Last Exam scores by cost (log scale), at each effort level. CursorBench 3.2.0 scores by cost (log scale), at each effort level. CursorBench 3.2.0 Accuracy vs Cost Fable 5.1 Fable 5 CursorBench 3.2.0 by cost (log scale), at each effort level. Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash in its internal systems that none of its engineers (or any other model) had been able to explain after several years of trying. Here, you can see how Fable 5.1 compares across various benchmarks: Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol Agentic scientific research Terminal-Bench-Science 0.1 [1] Agentic scientific research Terminal-Bench-Science 0.1 [1] 52.6% 24.7% 29.0% 22.4% Agentic coding Terminal-Bench 4.0 Agentic coding Terminal-Bench 4.0 55.8% 60.9% (Mythos 5.1) 42.0% 52.3% 37.3% Knowledge work GDPval-AA v2 Knowledge work GDPval-AA v2 1853 1723 1824 1711 Computer use OSWorld 2.0 [2] Computer use OSWorld 2.0 [2] 77.9% partial 72.9% partial 75.4% partial — partial Computer use OSWorld 2.0 Computer use OSWorld 2.0 41.7% strict 36.1% strict 39.6% strict — strict Multidisciplinary reasoning Humanity's Last Exam Multidisciplinary reasoning Humanity's Last Exam 60.9% no tools 57.8% no tools 56.6% no tools — no tools 65.0% with tools 63.8% with tools 63.6% with tools — with tools Business workflows AutomationBench Business workflows AutomationBench 31.4% 17.1% 26.9% 19.6% Agentic coding CursorBench 3.2.0 Agentic coding CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2% Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks. Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us: Previous 1 of 22 Next Quote “In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.” Company Jane Street Capital Author Craig Falls, Head of Quantitative Research

FDE 判断

这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。