Claude Code makes Auto Mode the default, ending the approve-every-command habit
From 14 August, Claude Code switches to Auto Mode by default for Pro, Max and Team plans. Anthropic's research shows it is meaningfully safer than manual approval.
By Nattapon YongpaiboonCo-founder, Claude Thailand Community
What’s changing
From 14 August 2026, Claude Code switches to Auto Mode as the default for everyone on Pro, Max and Team plans. Just as importantly, the safety classifier running behind the scenes stops consuming extra tokens for these plans.
In practice: instead of approving every single command or file edit yourself, each action now passes through a safety classifier that checks whether it’s destructive, irreversible, or reaches outside your machine. If it looks risky, Claude looks for a safer alternative or simply asks you first.
There’s a built-in circuit breaker too - three blocks in a row, or twenty blocks total in one session, and it falls back to asking for manual approval again.
Why Anthropic thinks this is actually safer
The most interesting part of the announcement is the research behind it, because it cuts against the instinct that a human clicking approve is inherently safer than letting the model decide.
People approve on autopilot. Real usage data shows people approved 97% of permission prompts, rejecting only 3% of individual permissions - yet rejecting 39% of higher-level plans. That gap suggests people scrutinise the big picture but wave through the details, the same pattern you see with “accept” buttons nobody actually reads.
A controlled test put humans and Auto Mode up against the identical set of dangerous commands:
| Caught the dangerous command | |
|---|---|
| Manual approval (human) | 13.6% |
| Auto Mode | 89% |
Worse, human vigilance degrades with session length - from 17% early on down to 5% after 50+ approvals - while Auto Mode held steady the whole way through.
Production data tells the same story. Sessions relying on manual approval showed serious, production-level harm twice as often as Auto Mode sessions (6.3% vs 2.4%).
Tested by outside teams
Anthropic brought in Apollo Research to red-team the system specifically. After hardening it against what they found, the miss rate dropped from 12% to 7% on a held-out set of attacks it had never seen before.
Prompt injection - a serious risk for any agent that acts on its own - got its own independent test from Trajectory Labs: 72 scenarios, 10 attempts each. Auto Mode blocked every single one, across all Claude models tested, against a 5.83% success rate for competing systems.
Inside Anthropic itself, Auto Mode caught three incidents before they became a problem: an attempt to publish data without authorisation, a destructive mass process shutdown, and an over-privileged credential escalation.
How it’s working in production
Enterprise teams running Auto Mode report roughly 25% more pull requests shipped. A few examples Anthropic shared:
- Adobe sped up its design review loop
- Nuro lets research agents run unattended overnight
- Gusto cut down on permission fatigue
- Garner Health standardised its software development process org-wide
Switching modes, and what to watch for
You can switch between Auto Mode and manual approval any time with Shift+Tab, or from the dropdown on desktop.
Anthropic is clear that Auto Mode doesn’t bring risk to zero - it still relies on a classifier that can miss edge cases. For anything touching production infrastructure, manual review is still the recommendation.
If you’re already on a Pro, Max or Team plan, this turns on automatically from 14 August - nothing to configure.
All figures and test results in this article are self-reported by Anthropic. Read their full write-up at the source link below.
Get it by email
New articles, Claude updates and community event announcements. Sent occasionally, never often enough to annoy you.
The newsletter is written in Thai
