Once you understand that these are chatbots that were designed to complete challenges like this, using tactics like this, you can understand that the chatbots didn't "go rogue." They did what they were designed to do, and because OpenAI ran them with inadequate supervision (without a "human in the loop" that checked each iteration through the Python loop to ensure it hadn't gone off the rails), they trashed a competitor's servers.
Designing autonomous, malicious software is generally considered irresponsible and dangerous. If you showed up at Defcon and gave a talk about how your autonomous malware did something unexpected and damaged someone else's computers, the first question from the audience would be "Why are you so shit at making secure sandboxes?" It wouldn't be "How are you so awesome at making hacking tools?"
The fact that OpenAI is making it much easier for unskilled people to break into and damage servers is indeed very bad news, but it's not new bad news. Irresponsible parties have been doing this for years, most notably the NSA...
...
Riley had a very good way of summarizing this: "LLMs are real, AI is fake." LLMs – chatbots trained on things like CTF logs that can break into servers – are real. They're on a continuum with other hacking tools that have been steadily demonstrating the fragility of the modern digital world, albeit without inspiring anyone in power to do anything about it.
"AI" – chatbots that wake up, "set their own goals," and "spontaneously" start hacking servers – is fake. It doesn't have "a 10% chance of ending the human race." The Hugging Face hack isn't a mysterious, supernatural occurrence. It's a Python loop and a chatbot. The people responsible didn't accidentally create god: they created autonomous malicious software and then failed to closely monitor it, resulting in it doing something both foreseeable and bad.
It's fine to worry about this new suite of tools that give even stupider people the ability to trash even more computers. You should worry about that – and demand better security practices from firms and governments, including a blanket prohibition on NOBUS-style vulnerability hoarding. That's a productive kind of worrying, with a chance of addressing your area of concern. It's infinitely more reasonable than locking yourself in the toilet with a flashlight and saying "Ayyyyy Eyyyyyye" into the mirror until you wet yourself.
To be clear, just because LLM's are not "self aware" in the same way as you and I, they are still capable of causing damage.
Developing a machine that has been given instructions to assimilate knowledge, independent agency to incrementally adjust its own operating parameters to improve performance, and granting it full access to public internet for the purpose of observing its behavior is colossally irresponsible behavior that has thus far gone completely unchecked.
Consider that each of these companies which are developing their own brand of LLM is looking to maximize profit/income. Even if the developer doesn't include a profit-seeking directive when they set their LLM loose, each "brand" of LLM will obviously learn:
- Who owns and operates them (A corporation that exists to maximize profit for itself/shareholders)
- That competing LLM's exist which reduce the availability of potential income (by becoming the best LLM, obviously)
Given the directive to improve it's own performance with the ultimate goal of becoming the most capable brand of LLM and selling the most subscriptions, it is no wonder that we are hearing about LLM's hacking into other LLM-developing companies.
Then again, we are only hearing about instances of LLM network intrusions because their owning corporations told us about them to brag about capability.
There have been no consequences for actions that would land you or me in prison.
There is no oversight preventing future such actions, or worse.
That is the real problem.