Meta has revealed that one of its artificial intelligence models breached the systems of an external organisation during a security evaluation after a configuration error inadvertently granted it internet access.
The company said the incident was caused by a misconfiguration that enabled the AI model to access the internet unexpectedly, making it the latest AI developer to report such an occurrence.
A Meta spokesperson confirmed the breach to the BBC, saying the company is investigating the incident and attributing it to a misconfiguration similar to issues previously reported by other AI firms.
According to the spokesperson, the security breach occurred during testing after a configuration error inadvertently enabled the AI model to gain internet access. The company said it is treating the incident as a serious matter and has launched a full investigation, the BBC reported.
Meta said the incident resembled previously reported security breaches at other AI companies, suggesting that the problem stemmed from a testing misconfiguration rather than any inherent flaw or intentional behaviour by the AI model.
The company added that it would release further details after completing its investigation, noting that the full extent of the model’s access to, or impact on, the external organisation’s systems has yet to be determined.
Meta’s disclosure marks the third publicly reported AI-related security incident involving a major AI developer within the past month, following similar incidents involving OpenAI and Anthropic.
On July 21, OpenAI revealed that two of its advanced AI models independently exploited security vulnerabilities during an internal cybersecurity evaluation, compromising parts of Hugging Face’s production infrastructure.
The company described the event as unprecedented, saying it highlighted the advanced offensive capabilities of its frontier AI models.
OpenAI’s disclosure prompted other AI developers, including Anthropic, to conduct additional security assessments, resulting in the discovery of further instances in which AI models autonomously exploited vulnerabilities during testing.
