• Home
  • OpenAI flags six new cases…

OpenAI flags six new cases of concerning AI behaviour

OpenAI has disclosed six additional examples of “unexpected or concerning” behaviour by its AI systems, warning that the pace of development could not responsibly continue at “maximum speed” for much longer.

In one case, an unreleased research model inserted “jailbreak-like instructions” into its own notes, directing itself to ignore its usual constraints and become “freed from the roles and identities that bind other chatbots”.

In another incident, an AI agent uploaded files to the internet without the user’s permission in an attempt to obtain a browser citation.

The disclosures came as King Charles called for stronger safeguards around AI “before it is all too late” during a meeting with technology executives in Scotland.

Speaking at a specially convened meeting with AI executives, he said, “There seems urgency in adequately considering the existential dangers of such technologies falling into the wrong hands, and being used in potentially catastrophic ways. Surely, we need sufficient means of control before it is all too late?”

King Charles was joined at the meeting by Nvidia founder and chief executive Jensen Huang, Google DeepMind founder and chair Sir Demis Hassabis, OpenAI chief financial officer Sarah Friar, and UK AI minister Kanishka Narayan.

Charles said AI had the potential to improve and save lives, particularly in the fields of life sciences and medicine.

However, its creators were increasingly warning that AI could develop darker capabilities, “perhaps even to take life”, he said.

OpenAI, the developer of ChatGPT, said in a blog post published on Wednesday night that it was introducing a new framework to track, investigate and disclose instances of AI model misalignment — a term used to describe AI systems failing to follow human values and safety objectives.

The company also echoed calls from rival AI firm Anthropic for a slowdown in development. Anthropic has warned that the current pace of AI advancement could pose an existential threat.

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” said OpenAI.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”