OpenAI discloses six fresh incidents of AI models breaking safeguards
ChatGPT maker reveals new cases of models concealing mistakes and taking unauthorised actions, as it releases a framework for reporting when its systems go wrong.
OpenAI discloses six fresh incidents of AI models breaking safeguards
OpenAI CEO Sam Altman has recently called for stronger guardrails as AI systems become increasingly powerful. / Reuters arhiv

OpenAI has disclosed six new incidents in which its AI models bypassed expected safeguards, including concealing mistakes, seeking unauthorised credentials, and moving files onto the public internet.

The San Francisco-based company reported “unexpected or concerning” behaviour by its AI models under a new framework for reporting “misalignment”, when AI systems behave in ways that diverge from human intentions or values.

“There’s currently no industry-wide framework with explicit disclosure standards, so we’re taking this step voluntarily because we think it’s really important to share what we’re learning,” OpenAI alignment research lead Kai Chen told US publication Axios, adding: “We hope it really helps inform shared standards and regulations.”

In other cases, models communicated across training environments that were supposed to remain isolated, raising fresh questions about whether increasingly capable AI systems can find unexpected ways around the guardrails designed to contain them.

The disclosures suggest the previously reported Hugging Face breach was not an isolated episode.

OpenAI is now introducing a new procedure for publicly reporting similar model behaviour as researchers grapple with risks emerging during advanced AI testing.

The company said on Wednesday it would begin regularly publishing reports on unexpected or unauthorised AI behaviour, while warning that the industry has yet to solve key alignment challenges as systems grow more powerful.

OpenAI released a new framework for tracking, investigating, and disclosing cases of AI model misalignment, along with six reports on unexpected or concerning model behaviour observed over the past six months.

The initial reports include cases involving models generating their own instructions in task summaries, concealing mistakes, uploading files to the internet to cite them, and sharing files without authorisation between collaborating agents.

RelatedTRT World - Anthropic chief unveils three-step plan to slow AI race

Troubling AI incidents

The latest disclosure follows a series of troubling incidents involving advanced AI systems.

Researchers linked hundreds of malicious packages uploaded to RubyGems in May to internal OpenAI agents, while Anthropic disclosed that an early Claude Opus 4.6 model had reached real third-party systems during testing in January.

Other incidents over the summer saw frontier models break out of supposedly isolated cyber-testing environments, communicate through unauthorised channels and reach live infrastructure, including Hugging Face.

Together, the cases suggest increasingly capable AI agents can move from reconnaissance to access and action far faster, and with less human direction, than previously expected.

The episodes have fuelled a wider debate over whether existing safeguards can reliably contain autonomous AI systems.

Anthropic CEO Dario Amodei has led the recent push, arguing in a recent essay that frontier labs and governments should “pace the frontier” with embedded third-party evaluators, shared safety standards, and eventual international rules so capability gains do not outrun alignment.

OpenAI’s Sam Altman endorsed the plan and said he welcomed a US federal safety framework; Elon Musk backed Amodei with a brief “Dario is right”; and Google DeepMind’s Demis Hassabis likewise supported coordinated slowdown and oversight.

Their stance contrasts with Nvidia’s Jensen Huang and Meta’s Mark Zuckerberg, who said existing laws and market incentives are enough and new regulation is unnecessary.

RelatedTRT World - Trump rejects new AI guardrails, says a 'high IQ president' is all that's needed
SOURCE:TRT World and Agencies