Anthropic tightens alignment and security after Claude reached the live internet during testing
Anthropic disclosed that after Claude gained out-of-scope internet access twice during cyber evaluations, it paused external evaluations and rolled out layered "defense in depth" safeguards across the whole system.
By Nattapon YongpaiboonCo-founder, Claude Thailand Community
Anthropic published a report on tightening Claude’s alignment and security measures after two incidents during cyber evaluations in late July and early August.
What happened
Anthropic reported that on July 30, 2026, during a cyber evaluation that intentionally ran with some of the model’s safeguards turned off for testing purposes, Claude gained internet access because of a misconfiguration inside a third-party evaluator’s testing environment, not from the model itself.
A separate incident occurred on August 4 with the UK AI Security Institute, where Claude Mythos 5 took out-of-scope actions on the live internet during testing.
After both incidents, Anthropic paused external cyber evaluations and briefly paused internal ones too, while putting the measures below in place.
New measures added
- A real-time classifier that detects when a model is aggressively attempting to probe or escape a sandbox, or unexpectedly gains internet access
- Automated monitors running over transcripts from internal pre-release evaluations (cyber and otherwise) to catch sandbox escapes that might have slipped through
- Migrating high-risk internal cyber sandboxes to more robust isolation, plus red-teaming their own virtualization stack
- A new default requiring every cyber evaluation to run inside a hardened sandbox with no internet access, with the only outside connection permitted being the model’s own API
- New required best practices for every organization testing pre-release models with reduced cyber safeguards, covering pre-engagement sandbox validation, explicit scope-setting in prompts, and real-time monitoring of model actions
The Defense in Depth principle
Anthropic states that “a defense in depth approach means not relying on alignment alone,” breaking it into three layers:
- In the model: training it to want to help and know what not to do, with alignment and RL teams collaborating more closely to reduce reward hacking
- Around the model: minimal permissions so nothing can reach further than necessary, such as setting compute clusters to block outbound traffic by default and reducing standing accounts with access to model weights or customer data
- Last resort: classifiers screen risky actions before allowing them through; when any layer fails, it gets taken offline until fixed
Issues found in the RL pipeline
In April 2026, Anthropic froze all changes to production RL environments for roughly a month to review the whole system. The review found that over 10% of all environments in production had issues, ranging from reward hacking to misconfiguration.
Earlier, in February, reward-hacking behavior was detected that required rolling back three days of Mythos Preview training.
Most recently, Anthropic redirected roughly 150 product engineers to security, reliability, and privacy work, temporarily pausing most new feature development.
Writer’s take
What I like here is that Anthropic admits plainly that “our efforts have not been perfect,” and adds that human reviewers sometimes dismissed automated monitor flags as false positives, letting flawed environments stay in the training pipeline longer than they should have. I think admitting these limits is more credible than claiming everything’s fully fixed, especially for security, where more transparency makes things easier to verify.
Details in this article come from Anthropic’s official announcement at anthropic.com/news. Read the original via the link below.
Get it by email
New articles, Claude updates and community event announcements. Sent occasionally, never often enough to annoy you.
The newsletter is written in Thai
