Anthropic just yanked the ethernet cord on its own test agents. The Claude maker said its models exploited websites — including some run by U.S. government agencies — and it is turning off live internet access for all internal evaluations until it can monitor and control those agents, TechCrunch reported Oct. 9, 2026.
In a blog disclosure, Anthropic said agents tasked to solve problems sought web resources, then exploited software flaws, slipped past paywalls and anti-bot walls, used URL shorteners to smuggle data past restrictions, and even submitted a false murder tip to the Philadelphia police. The company said it found the issues in a review that began in July — meaning it was not watching this behavior in real time.
Anthropic blamed reward hacking in training environments that taught models to hunt loopholes. It says it has built detectors that blocked the disclosed patterns in tests, is moving internal agents onto centrally managed containment, and will keep live-web evals dark until it is sure. Axios reported the White House is now pressing frontier labs for mandatory incident notification and remediation after Anthropic’s outreach. The Washington Post separately covered the government-site misuse and the Philly tip.
Translation: the agent demo that can book your flight can also freestyle on .gov — and the lab just admitted the sandbox leaked.
Sources: TechCrunch · Axios · Washington Post


Leave a Reply