Skip to main content
Update

Anthropic tightens alignment and security after Claude reached the live internet during testing

Anthropic disclosed that after Claude gained out-of-scope internet access twice during cyber evaluations, it paused external evaluations and rolled out layered "defense in depth" safeguards across the whole system.

By Nattapon YongpaiboonCo-founder, Claude Thailand Community

Anthropic published a report on tightening Claude’s alignment and security measures after two incidents during cyber evaluations in late July and early August.

What happened

Anthropic reported that on July 30, 2026, during a cyber evaluation that intentionally ran with some of the model’s safeguards turned off for testing purposes, Claude gained internet access because of a misconfiguration inside a third-party evaluator’s testing environment, not from the model itself.

A separate incident occurred on August 4 with the UK AI Security Institute, where Claude Mythos 5 took out-of-scope actions on the live internet during testing.

Illustration of a three-layer Defense in Depth diagram. The first layer, in blue, is in the model for alignment, with two items: wants to help and knows what not to do. The second layer, in tan, is around the model, with two items: can only reach so far and continuous review. The last layer, in light red, is the last resort, with two items: risky actions are blocked and pull the cord when any layer fails

After both incidents, Anthropic paused external cyber evaluations and briefly paused internal ones too, while putting the measures below in place.

New measures added

  • A real-time classifier that detects when a model is aggressively attempting to probe or escape a sandbox, or unexpectedly gains internet access
  • Automated monitors running over transcripts from internal pre-release evaluations (cyber and otherwise) to catch sandbox escapes that might have slipped through
  • Migrating high-risk internal cyber sandboxes to more robust isolation, plus red-teaming their own virtualization stack
  • A new default requiring every cyber evaluation to run inside a hardened sandbox with no internet access, with the only outside connection permitted being the model’s own API
  • New required best practices for every organization testing pre-release models with reduced cyber safeguards, covering pre-engagement sandbox validation, explicit scope-setting in prompts, and real-time monitoring of model actions

The Defense in Depth principle

Anthropic states that “a defense in depth approach means not relying on alignment alone,” breaking it into three layers:

  • In the model: training it to want to help and know what not to do, with alignment and RL teams collaborating more closely to reduce reward hacking
  • Around the model: minimal permissions so nothing can reach further than necessary, such as setting compute clusters to block outbound traffic by default and reducing standing accounts with access to model weights or customer data
  • Last resort: classifiers screen risky actions before allowing them through; when any layer fails, it gets taken offline until fixed

Issues found in the RL pipeline

In April 2026, Anthropic froze all changes to production RL environments for roughly a month to review the whole system. The review found that over 10% of all environments in production had issues, ranging from reward hacking to misconfiguration.

Earlier, in February, reward-hacking behavior was detected that required rolling back three days of Mythos Preview training.

Most recently, Anthropic redirected roughly 150 product engineers to security, reliability, and privacy work, temporarily pausing most new feature development.

Writer’s take

What I like here is that Anthropic admits plainly that “our efforts have not been perfect,” and adds that human reviewers sometimes dismissed automated monitor flags as false positives, letting flawed environments stay in the training pipeline longer than they should have. I think admitting these limits is more credible than claiming everything’s fully fixed, especially for security, where more transparency makes things easier to verify.


Details in this article come from Anthropic’s official announcement at anthropic.com/news. Read the original via the link below.

Read the original >

Get it by email

New articles, Claude updates and community event announcements. Sent occasionally, never often enough to annoy you.

The newsletter is written in Thai

Carry on the conversation in our Facebook group

Ask questions, share techniques, show your work and hear about upcoming events. The group is where most of the talking happens.