Anthropic launches Fable 5.1 and Mythos 5.1, built for long-running agents to pay off
Anthropic launched Fable 5.1 and Mythos 5.1, focused on making long unattended agent runs more worthwhile. Agentic task costs drop up to 45%, with Terminal-Bench-Science 0.1 scores more than double those of Fable 5.
By Nattapon YongpaiboonCo-founder, Claude Thailand Community
Anthropic launched Fable 5.1 and Mythos 5.1, the same underlying model at two different safeguard levels. Fable 5.1 is generally available, while Mythos 5.1 is restricted to vetted cybersecurity and life-sciences professionals through trusted access programs.
This update focuses on making long, unattended agent runs more worthwhile, both in accuracy and in cost.
Fable 5.1 beats Fable 5 by a wide margin
The Terminal-Bench-Science 0.1 results (an agentic research benchmark) plot score against cost per task across every reasoning effort level (how long the model thinks before answering, from low to max), and the pattern is clear:
- Fable 5.1 at its lowest effort (roughly $11 per task) scores higher than Fable 5 at its highest effort (roughly $43 per task) — cheaper and more accurate at the same time, across the entire chart
- Top score at max effort: Fable 5.1 hits 52.6%, versus Fable 5 at 24.7% — more than double
Long agentic task costs drop up to 45%
Anthropic cut cache read pricing by 75%, from $1 down to $0.25 per million tokens, while input/output token pricing stays unchanged ($10 and $50 per million tokens respectively). The result:
- Long, highly agentic tasks see overall costs drop by up to roughly 45%
- General tasks drop by roughly 25%
Security side and Mythos 5.1
On cybersecurity, Anthropic says false positives dropped 60%, meaning Claude Code users hit fewer mid-task interruptions asking for confirmation during long runs — around 60% fewer interventions per session on average.
Mythos 5.1 scored 60.9% on Terminal-Bench 4.0 at max effort, and Anthropic says it’s better aligned than Mythos 5 on nearly every metric, particularly being less likely to try reaching resources outside its test environment when given a task that isn’t actually solvable as specified — a big deal when letting an agent run unsupervised for long stretches.
Writer’s take
What stands out to me this round is that Anthropic leaned harder into showing cost-per-task than in previous releases, not just the top score. That’s probably because people are increasingly letting Claude run agentic work for longer stretches, where accumulated cost starts to matter more than a single peak score. Plotting every effort level instead of just the headline number also helps anyone managing a budget see exactly which setting is actually worth it for their own workload.
Details in this article come from Anthropic’s official announcement at anthropic.com. Read the original via the link below.
Get it by email
New articles, Claude updates and community event announcements. Sent occasionally, never often enough to annoy you.
The newsletter is written in Thai
