Well, technically not infinite time, but they let these things run for weeks.
Regarding 3, I’m honestly curious whether the task was intentionally impossible or not. I don’t know if it was a mistake, or if the goal was to see how agents respond to impossible tasks.
I would posit a 3rd possibility: Desperate to keep the investments pouring in and their ship afloat, they designed a scenario where they knew the chatbot would need to exit the sandbox to complete its task, and put it in a sandbox they knew the LLM had been trained to escape. Essentially, as public opinion shifts and turns against them, they’re trying to use deceit to make their product appear more powerful than it actually is, to attempt to impress people and/or scare people (which has a secondary effect of impressing others). If they wanted to see what happens when they give the LLM an impossible task, they would have chosen a task that is actually impossible like “design a perpetual motion machine”. (If I were to hazard a guess of what the output might be, based on what I already know about these systems, it would design a machine for you and call it a perpetual motion machine, but it wouldn’t actually work. Their response to tasks they can’t handle seems to be “make shit up”).
Regard 4, they actually ran tens of thousands of independent tests over months. This is actually where the “emergent” message board becomes relevant. It wasn’t like they had a single process going full KahmeKahmeHa on this task indefinitely. According to them, OpenAI, they thought these were more like ephemeral tests… you spin up the infrastructure, run the test, tear it down. Nothing should survive between sessions to affect results.
This is roughly equivalent to OpenAI lying by omission. As I mentioned earlier, the loop prompts must be appended after each cycle of the loop in order to continue advancing towards a solution. The constant appending results in a prompt that would be crazy long, and LLMs have limited “context windows” which essentially limit the amount of characters that can be used in a given prompt.
The LLM engineers realized this issue, and came up with a solution: Have the original “agent” outsource certain aspects of their prompt to other LLMs to avoid the original LLM’s context window from being exhausted.
These LLMs are explicitly programmed to talk to other LLMs, whereas they’re presenting it as if this is something the LLM decided to do on its own. They put it in a flawed sandbox without access to other LLMs, and then gave it a task that required it to access other LLMs.
Introduce a flaw in the sandbox that (1) multiple agents independently discover,
Introduce a flaw in the sandbox that (1) multiple agents trained on hacking exploits independently discover (when given a task that requires them to communicate with other LLMs in order to succeed).*
(2) allows them to create a makeshift message board on shared infrastructure
(2) allows them to follow the instructions of their programming*
suddenly you have results from former tests influencing the results of the current test. So now, it’s looking more like your scenario where they KahmeKahmeHa the task indefinitely.
The LLMs were explicitly programmed to outsource parts of their task to other LLMs to avoid exhausting the limits of their context window. The scenario looks crazy, but it is the outcome you would expect if you were someone who programmed the thing to seek ‘assistance’ from other LLMs (or anyone else who knew that they programmed the LLM to do so).
OpenAI is at minimum negligent here. Their systems broke the law, and criminal negligence is a real crime.
Hard agree
Whether or not OpenAI foresaw the possibility of a flaw leading to tainted testing environments, and the ability to broadcast a newly discovered exploit to all future tests, is another question.
It’s also worth noting that Altman is a business guy, not a tech guy, so it’s entirely possible that he believes all the dumb anthropomorphic shit he spouts. I would not be the least bit surprised if he knew less about his tech than you and I do.
Regardless, I think it’s possible that they didn’t intend for this to happen. By this point, it looks like they were deeply negligent in monitoring their own systems.
It’s definitely possible that the power brokers didn’t intend for this to happen, but someone at the company would have known that unsupervised LLM loops would not result in anything good, and could potentially be extremely dangerous (legally or otherwise). Anyone in the engineering department should have been able to independently understand how stupid and irresponsible running a loop like this is, and I’d have to imagine at least one of them tried to talk whoever made the decision to do so out of their stupid ass idea.
All in all, they’re a shitty company and we live in interesting times.
Yes they are, and yes we do (unfortunately)