• Home
  • OpenAI halts some Astra activities…

OpenAI halts some Astra activities to strengthen AI safeguards

OpenAI has paused some internal work on an upcoming artificial intelligence model to strengthen its safeguards after discovering that the system was significantly more capable of performing cybersecurity tasks.

The ChatGPT maker said on Friday that it “cannot rule out” the unreleased Astra model reaching OpenAI’s “critical cybersecurity threshold,” meaning it could potentially identify and develop zero-day exploits without human intervention.

OpenAI said it is taking steps to strengthen security controls for the development and testing of new AI models, while pausing internal activities involving Astra that do not meet the enhanced requirements.

The move comes amid growing concerns over the cybersecurity capabilities of advanced AI models.

Over the past two weeks, OpenAI and Anthropic have acknowledged that their models inadvertently breached the systems of multiple organisations, including Hugging Face, while being tested.

Meta also said on Wednesday that a recently released AI model had infiltrated a third party’s computer system.

In its blog post, OpenAI said it would work with government agencies and AI safety organisations to test Astra’s capabilities.

The company also plans to provide guidance to third-party testing partners on how to safely evaluate its more advanced AI models.