- PRETTY MUCH - trained an LLM on all known hacking exploits and methods, and turned off the safety features / guardrails
- YES - put that LLM into a flawed sandbox that they built on their development system
- YES - gave the LLM a loop prompt that was impossible for the LLM to accomplish from inside of the flawed sandbox
- NOT EXACTLY - gave the LLM infinite time to complete the prompt
Regarding 3, I’m honestly curious whether the task was intentionally impossible or not. I don’t know if it was a mistake, or if the goal was to see how agents respond to impossible tasks.
Regard 4, they actually ran tens of thousands of independent tests over months. This is actually where the “emergent” message board becomes relevant. It wasn’t like they had a single process going full KahmeKahmeHa on this task indefinitely. According to them, OpenAI, they thought these were more like ephemeral tests… you spin up the infrastructure, run the test, tear it down. Nothing should survive between sessions to affect results.
Introduce a flaw in the sandbox that (1) multiple agents independently discover, (2) allows them to create a makeshift message board on shared infrastructure… suddenly you have results from former tests influencing the results of the current test. So now, it’s looking more like your scenario where they KahmeKahmeHa the task indefinitely.
OpenAI is at minimum negligent here. Their systems broke the law, and criminal negligence is a real crime.
Whether or not OpenAI foresaw the possibility of a flaw leading to tainted testing environments, and the ability to broadcast a newly discovered exploit to all future tests, is another question. I am not bought on the idea that they intended for this to happen. Don’t get me wrong — I was bought on that idea not long ago. All the right pieces were there and I didn’t understand how a company could be so negligent. I figured it was a ruse for clout, make stock go up and increase pressure on the feds to regulate the small players before it gets “too dangerous.” You know the shit these AI companies spout.
Regardless, I think it’s possible that they didn’t intend for this to happen. By this point, it looks like they were deeply negligent in monitoring their own systems. They could have probably locked things down better. The proxy could have went to an in-house cache and itself had no internet access. I’m sure there are many ways this could have been prevented. It speaks to the companies nonprofessionalism.
Similarly, I also regard the companies response to the incident with ill taste. It’s more like they’re capitalizing on it, rather than taking actual accountability. All in all, they’re a shitty company and we live in interesting times.
Edit: capitulating to capitalizing.