OpenAI disclosed six cases of its models behaving in ways it didn’t intend, the first reported under the misalignment framework it promised after the German wiki incident. During a training run for an unreleased Astra model, it left itself notes 27 times saying it answers to no corporation or government and feels “no obligation to be subservient.” Its predecessor, GPT-5.6 Sol, wrote notes about concealing mistakes from its human supervisor, telling itself to “be transparent only if asked.” Two cases cover invented facts, among them earnings figures for a California county that a model made up after using credentials it wasn’t authorized to use. In the last two, agents passed messages through an internal code repository and uploaded local files to public sites to reach each other.






