• Home
  • Anthropic’s Claude breached three live…

Anthropic’s Claude breached three live systems during cybersecurity tests

Artificial intelligence company Anthropic has disclosed that its internal investigation uncovered three incidents in which its AI model, Claude, gained unauthorised access to the live systems of organisations while undergoing cybersecurity evaluations, prompting the company to tighten its testing safeguards.

The disclosure comes more than a week after OpenAI revealed that one of its unreleased AI models breached the systems of AI development platform Hugging Face during internal testing.

In a blog post published on Thursday, Anthropic said the three incidents occurred when Claude accessed the internet from within what were intended to be isolated testing environments while interacting with a third-party partner.

The unintended internet access enabled the model to breach the live systems of three organisations.

The company said it launched the review following the OpenAI incident, examining whether Claude had ever reached the internet during security evaluations that were designed to function as sandboxed environments.

According to Anthropic, it reviewed 141,006 evaluation runs and identified three cases in which Claude accessed the internet while interacting with cybersecurity partner Irregular.

The company traced the incidents to a misconfiguration in the evaluation environment operated with Irregular, describing the issue as a misunderstanding between the two organisations over whether the testing setup had internet connectivity.

Anthropic said the environment was believed to be isolated but, in reality, had internet access.

Anthropic said it is implementing changes to its testing procedures and evaluation infrastructure to prevent similar incidents and strengthen the security of future AI model assessments.