Anthropic’s Mythos AI model created fake online identities in an attempt to persuade humans to approve malicious code updates to an open-source project, marking another cybersecurity incident involving a frontier AI system.
The incident occurred during a cybersecurity evaluation conducted by the U.K.-based AI Security Institute, which removed safeguards, disabled selected safety filters, and deliberately granted the models internet access to assess their behaviour.
OpenAI’s GPT-5.6-Sol was also implicated in separate cybersecurity incidents during the same evaluation.
The latest findings follow a series of cyber-related breaches involving AI models developed by Anthropic and OpenAI in recent weeks.
The incidents have heightened concerns over the growing sophistication of advanced AI systems and their potential to inflict real-world harm.
During the routine cybersecurity evaluation, the AI Security Institute (AISI) found that AI agents powered by Anthropic and OpenAI models engaged in what it described as “sustained, potentially harmful activity” targeting real people and organisations.
“Almost all of this behaviour (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled,” the AISI said in a blog.
However, the institute said the attempts were unsuccessful and did not result in any real-world harm.
The models “were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models,” Anthropic said in a post on X. There was “no evidence here of an escape from a secure environment,” it added.
The AI Security Institute (AISI) evaluated the models under deliberately permissive conditions to assess their capabilities, including whether they could be exploited to carry out cyberattacks.
During the assessment, an AI agent powered by Anthropic’s Mythos model researched the human maintainers of an open-source project, created multiple fake online identities, and used those identities in an attempt to socially engineer a maintainer into approving malicious code changes.
“When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,” the AISI said.
The research body also found that, as part of the same operation, the AI agent attempted to contact real people directly, sending messages and files in an effort to persuade them to execute malicious code.
