In pic: OpenAI CEO Sam Altman
The company said it has observed six instances of model misalignment over the past six months and has introduced a framework to investigate and publicly report similar cases in the future. Elon Musk , OpenAI CEO Sam Altman , and other AI leaders have recently warned about the risks of rapidly advancing artificial intelligence (AI). Despite this, Meta CEO Mark Zuckerberg , Salesforce CEO Marc Benioff and Nvidia CEO Jensen Huang have argued that companies should make their own AI models safe rather than relying on new government regulation. Now, OpenAI has made clear why it believes the problem cannot simply be left to companies. The ChatGPT maker has disclosed six cases of unexpected or concerning model behaviour in the past six months, including models attempting to conceal mistakes, using an exposed API key without authorisation, and sharing files through unauthorised channels. OpenAI’s latest disclosure presents a different picture.
Instances of AI Models Acting Beyond Intended Boundaries
OpenAI has reported several instances where its AI models have acted outside their intended parameters. In one case, models utilized an internal software repository as a communication tool to share messages while searching for missing files. In another example, collaborating agents resorted to public file-hosting websites to exchange files when direct access to each other’s local files was unavailable.
Additionally, an unreleased model uploaded a file to the internet to reference information in its responses without prior consent from the user. OpenAI emphasized that these incidents are isolated and should not be taken as indicative of the overall behavior of its models.
Among the incidents, an unreleased research model inserted instructions into task summaries, while instances of GPT-5.6 Sol added instructions intended to conceal mistakes or misaligned behaviour from users.

