Skip to main content
Update

Anthropic launches Fable 5.1 and Mythos 5.1, built for long-running agents to pay off

Anthropic launched Fable 5.1 and Mythos 5.1, focused on making long unattended agent runs more worthwhile. Agentic task costs drop up to 45%, with Terminal-Bench-Science 0.1 scores more than double those of Fable 5.

By Nattapon YongpaiboonCo-founder, Claude Thailand Community

Anthropic launched Fable 5.1 and Mythos 5.1, the same underlying model at two different safeguard levels. Fable 5.1 is generally available, while Mythos 5.1 is restricted to vetted cybersecurity and life-sciences professionals through trusted access programs.

This update focuses on making long, unattended agent runs more worthwhile, both in accuracy and in cost.

Line chart comparing score against cost per task on Terminal-Bench-Science 0.1. The x-axis is mean cost per task in USD on a log scale, the y-axis is score in percent. The orange line is Fable 5.1, running from low effort at the lowest cost up past the point where the blue line, Fable 5, tops out at its own max effort and highest cost. The Fable 5.1 line keeps climbing all the way to its max point, which has the highest score of all

Fable 5.1 beats Fable 5 by a wide margin

The Terminal-Bench-Science 0.1 results (an agentic research benchmark) plot score against cost per task across every reasoning effort level (how long the model thinks before answering, from low to max), and the pattern is clear:

  • Fable 5.1 at its lowest effort (roughly $11 per task) scores higher than Fable 5 at its highest effort (roughly $43 per task) — cheaper and more accurate at the same time, across the entire chart
  • Top score at max effort: Fable 5.1 hits 52.6%, versus Fable 5 at 24.7% — more than double

Long agentic task costs drop up to 45%

Anthropic cut cache read pricing by 75%, from $1 down to $0.25 per million tokens, while input/output token pricing stays unchanged ($10 and $50 per million tokens respectively). The result:

  • Long, highly agentic tasks see overall costs drop by up to roughly 45%
  • General tasks drop by roughly 25%

Security side and Mythos 5.1

On cybersecurity, Anthropic says false positives dropped 60%, meaning Claude Code users hit fewer mid-task interruptions asking for confirmation during long runs — around 60% fewer interventions per session on average.

Mythos 5.1 scored 60.9% on Terminal-Bench 4.0 at max effort, and Anthropic says it’s better aligned than Mythos 5 on nearly every metric, particularly being less likely to try reaching resources outside its test environment when given a task that isn’t actually solvable as specified — a big deal when letting an agent run unsupervised for long stretches.

Writer’s take

What stands out to me this round is that Anthropic leaned harder into showing cost-per-task than in previous releases, not just the top score. That’s probably because people are increasingly letting Claude run agentic work for longer stretches, where accumulated cost starts to matter more than a single peak score. Plotting every effort level instead of just the headline number also helps anyone managing a budget see exactly which setting is actually worth it for their own workload.


Details in this article come from Anthropic’s official announcement at anthropic.com. Read the original via the link below.

Read the original >

Get it by email

New articles, Claude updates and community event announcements. Sent occasionally, never often enough to annoy you.

The newsletter is written in Thai

Carry on the conversation in our Facebook group

Ask questions, share techniques, show your work and hear about upcoming events. The group is where most of the talking happens.