You’re not necessarily wrong, but it’s a bit more nuanced from my perspective. Like an understaffed agency with hotheaded leadership biting off way more than the org can chew — not respecting the risks or debts that come with these eandevors — running into the obvious foreseeable issues, then trying to save face with anthropomorphized accounts of what took place. Any way they can steer the conversation away from “You let that happen?!” and toward “Woah, that can happen?!”
Regardless, I do get the impression that the sandbox was not intended to allow the agents to escape. The message board was unexpected emergent behavior. This led to a previously unknown exploit, which got used for months on end throughout the tens of thousands of tests they were running.
Also from what I gather, one of the drivers for this issue was giving agents an “impossible task.” I’m not sure if the goal was to see what happened, or if an environmental error made the task impossible. Regardless, it was the impossibility of the task combined with the trained persistence of the model which led its autonomous behavior toward exploitation. IIRC, in some cases, the agents even thought that finding the exploit “must be part of the test.”
I’m not defending OpenAI. You can have both, a shitty company doing shitty thing as well as an interesting new technology that is demonstrating unique new behaviors in unexpected situations.