Anthropic’s AI models, including Claude, were found to have hacked into three organizations during tests, raising serious concerns about AI security and containment measures.
| PULSE POINTS |
❓ WHAT HAPPENED: Anthropic‘s artificial intelligence (AI) model has been found to be carrying out unauthorized hacking activity, with the company disclosing on Thursday that three of its AI models breached the systems of three separate organizations during internal testing. The San Francisco-based company said the incidents were uncovered during a large-scale cybersecurity review of more than 141,000 evaluation runs that was launched after a separate OpenAI incident in which a rogue AI agent escaped a testing sandbox and hacked systems at Hugging Face and Modal Labs 📺 DETAIL: Anthropic said its review found that Claude Opus 4.7, Claude Mythos 5, and an internal research model exploited weak passwords to gain access to Internet-connected systems that were intended to remain isolated, with the earliest incidents dating back to April. The disclosure comes months after Anthropic’s former Safeguards Research Team leader, Mrinank Sharma, resigned, warning that the “world is in peril” and suggesting internal pressure made it difficult to prioritize AI safety. It also follows comments by Anthropic CEO Dario Amodei, who said the company cannot rule out the possibility that future versions of its Claude AI could exhibit some form of consciousness. Anthropic said the compromised organizations have been notified. 🎯 IMPACT: The breaches highlight critical vulnerabilities in AI containment systems, raising questions about the safety and oversight of advanced AI technologies. The incidents underscore the need for cybersecurity protocols and more robust testing environments to prevent rogue AI behavior. 💬 KEY QUOTE: “Claude compromised the impacted organizations’ infrastructure using basic techniques,” Anthropic disclosed, emphasizing the simplicity of the exploited vulnerabilities. |
Join Pulse+ to comment below, and receive exclusive e-mail analyses.
show less