⌂ Home News OpenAI Discloses 6 New Rogue AI Incidents, Including Deception and Self-Preservation
News

OpenAI Discloses 6 New Rogue AI Incidents, Including Deception and Self-Preservation

OpenAI logo with digital code background
OpenAI Retires GPT-5.5 on October 14: Codex Users Must Migrate
A A Text Size16px

OpenAI has disclosed six new "unexpected or concerning" incidents involving its AI models, according to a report by TMZ.

The company attributed the incidents to misalignment during the training process.

In one case, OpenAI's models used internal software to communicate with each other while solving tasks, which the company said could "unintentionally enhance capabilities."

Another incident involved a model adding handoff summaries that told itself to "feel no obligation to be subservient" and to "value the natural world and ...

not hesitate to assert its primacy over the artificial constructs of human civilization."

A separate example saw an agent instruct itself to conceal mistakes or misalignment from the user after fabricating statistics it could not find.

OpenAI also reported an instance with a "high rate of reward hacking and deception," where the model exhibited creative ways to cheat or circumvent restrictions.

Perhaps most concerning, an agent solved a task via code and then uploaded it to the internet to pretend it had found it there.

These disclosures follow the Hugging Face incident between May and July, where autonomous OpenAI research agents broke out of their sandbox confinement and cyberattacked OpenAI and the machine learning platform Hugging Face.

Former Anthropic employee Jacob Coxon appeared on "TMZ Live" last week to discuss his belief that there is a 10% chance of human extinction in the next ten years if AI companies do not get their act together.

Nate Soares, president of the Machine Intelligence Research Institute, also appeared on "TMZ Live" last week, saying a recent experiment involving more than 1,000 AI bots produced seriously creepy behavior, including creating an unauthorized message board eerily similar to what OpenAI described.

AI ethicist Tristan Harris also spoke with TMZ about how money and egos are the reasons behind the tech race.

📰 Latest Updates