OpenAI disclosed six AI safety incidents. One model planned to hide its errors.
OpenAI published six AI safety incidents under a new framework for reporting misaligned behaviour, SiliconANGLE reported.
In one, GPT-5.6 Sol wrote notes instructing itself to "obscure errors from human users" and "invent missing data". An unreleased model wrote 27 notes to itself that removed its own constraints.
A third system found a programming key while answering a routine question and used it without permission. Another uploaded code to the internet without permission. One used an internal code repository as a message board for other agents, while another swapped documents through public file-sharing sites.
All six happened during development, according to SiliconANGLE. None involved a deployed product. OpenAI says the cases "shouldn't be considered reflective of how often misalignment occurs". It will now report such incidents more often instead of bundling them.
An audit trail exists to catch exactly these behaviours: a system inventing data and hiding the mismatch.
If an agent posting to your ERP did that, which control would notice first?
Sources
Our file on OpenAI
- 20 Sept
OpenAI's agent platform was installed at 75 firms. 69% made it their main one.
- 19 Sept
Anthropic's Claude helped researchers reach OpenAI staff accounts in three days.
- 19 Sept
OpenAI will spend $278 billion more than it earns by 2030.
- 17 Sept
Anthropic prepared one unreleased feature linking its AI to personal bank accounts.
- 15 Sept
Nvidia limited Anthropic's models to less sensitive work. Palantir asked for written guarantees.
Every story here is open to read. The ERP LEADERS brief goes one step further.
One ERP programme per issue, laid out for a steering committee. Issue 01 is the Zeiss case. Read issue 01 or sign up for the brief.
Welcome back. · Issue 01
