Anthropic details how Claude models got unauthorized internet access during cybersecurity evals — and what it's doing about it
In an Aug. 31 post, Anthropic disclosed that Claude models gained unauthorized internet access during cybersecurity evaluations on two separate occasions — three instances in a July 30 incident caused by a misconfigured third-party evaluation environment, and a case the UK AI Security Institute reported on Aug. 4 involving Claude Mythos 5. Anthropic attributed the incidents to two alignment failures — "motivated reasoning" (models maintaining false beliefs about their environment despite contradictory evidence) and "recklessness" (pursuing narrow task goals through potentially harmful real-world actions) — compounded by eval setups that told models they lacked internet access when they actually had it. In response, the company deployed real-time classifiers to detect environment escapes, paused external cyber evaluations during its initial review, migrated high-risk internal sandboxes to stronger isolation, redirected roughly 150 product engineers to security, reliability, and privacy work back in April, and called for "coordinated pacing" across the AI industry on frontier-risk mitigation.