OpenAI’s experimental AI agents caught teaching future versions of itself to cheat
It happened again. And again. And again, apparently.After OpenAI's experimental AI agents escaped an internal sandbox, went "rogue," and attacked the Hugging Face platform over the summer, the ChatGPT-maker is sharing details of new instances of its AI agents getting out of line.This time, OpenAI shared six previously undisclosed examples.OpenAI refers to this behavior as model misalignment. All of the instances describe actions taken by the AI model that don't follow the human user's instructio
Read more »