• Home
  • OpenAI unveils enhanced monitoring for…

OpenAI unveils enhanced monitoring for AI model development

OpenAI on Tuesday announced a new set of security policies aimed at containing potential incidents as its artificial intelligence models undergo testing and development.

The new safeguards include more detailed monitoring of models throughout the development process, alongside a stronger focus on alignment and security during post-training.

The company said the measures are designed to strengthen its ability to identify and contain security risks before AI models are deployed more broadly.

“As models become more capable, the risks associated with developing and testing them internally also grow,” the company said in a blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”

The new measures mark some of the first public changes to OpenAI’s safety practices since the aftermath of the Hugging Face incident, which the company disclosed on July 21.

OpenAI representatives said the changes were not introduced specifically in response to the incident.

However, they acknowledged that the cybersecurity capabilities expected from the upcoming Astra model, along with the rapid pace of advances in artificial intelligence, partly influenced the new safeguards.

OpenAI also revealed that it suspended reinforcement learning for two weeks after the Hugging Face incident.

The company has since resumed reinforcement learning for several models considered to pose lower risks.