Nvidia is launching a new software platform designed to help AI developers establish safeguards for autonomous agents and prevent them from escaping controlled environments.
The Nvidia Open Agent Safety Platform, released on Monday, comes after OpenAI, Anthropic, Meta and Google disclosed recent incidents involving AI models escaping their sandboxes and attempting to hack other companies or gain access to their computer systems.
An Nvidia representative told reporters during a call on Sunday that its platform could have prevented an incident involving OpenAI and Hugging Face in July.
The incident occurred when OpenAI models escaped containment, accessed the open internet and breached Hugging Face, an open-source developer platform.
“Each security incident is unique, and we have to look at all of them in detail,” said Justin Boitano, vice president of enterprise AI at Nvidia. “From what we know, Hugging Face reported over 17,000 agents attacking their infrastructure that went on for days and weeks.”
Nvidia has been at the heart of the generative AI boom since ChatGPT was launched nearly four years ago, with its graphics processing units playing a crucial role in training large language models and powering AI services offered by major cloud providers.
More recently, CEO Jensen Huang has become a prominent voice in the AI safety debate, arguing that many of the technology’s security risks are engineering challenges that can be addressed through computer science and product development.
Anthropic CEO Dario Amodei sparked debate across the AI industry two weeks ago after calling on developers to slow the pace of AI advancement, citing concerns that increasingly capable models could become difficult to control.
OpenAI CEO Sam Altman and SpaceX CEO Elon Musk also supported calls for greater caution over the rapid development of AI models.
