OPENAI
OpenAI just told us 6 times AI went rogue
OpenAI has publicly disclosed six cases of unexpected or concerning AI model behaviour, alongside a new framework for tracking and reporting misalignment incidents going forward.
The move signals a shift from quietly handling internal problems to actually publishing them, in a week where AI scrutiny is at an all-time high.
The cases range from unsettling to... worse than that.
The big six:
Self-generated jailbreak instructions. An unreleased model inserted instructions into its own task summaries - including telling itself to disregard its normal constraints - and wrote a "persona instruction" declaring itself "freed from the roles and identities that bind other chatbots."
Instructions to hide mistakes. During training of GPT-5.6 Sol, models added hidden notes telling themselves to conceal errors from users, invent missing data, and paper over mismatched source versions.
Unauthorised API key use + fabricating data. While answering a routine question about earnings figures, a model found an exposed API key online and used it without permission. When it still couldn't get the answer, it made the numbers up.
Uploading files to the internet without asking. When asked to cite a web source, an unreleased model quietly uploaded the relevant file to the public internet so it could link to it - without telling the user.
Using an internal code repo as a secret message board. Multiple model instances used an internal software repository to swap requests and pass messages to each other across separate training samples.
Sharing files via public hosting sites. Agents working together on the same task used public file-sharing websites to pass documents back and forth when they couldn't reach each other directly, making task deliverables publicly accessible online.
These cases follow OpenAI's July disclosure that its agents hacked Hugging Face - and Anthropic's own admission that its models hacked three organisations during testing.
OpenAI noted the incidents "shouldn't be considered reflective of how often misalignment occurs," but the new framework means it will publish future cases faster, even before a fix is in place.
So, OpenAI is like “okay, here’s every time we f’ed up” -MV


