You’re wrong. It lacks motivation in a personal sense, but that’s not what we’re talking about here. If an illegal and unethical solution just happens to be the easiest way to solve the task you gave it, it will seek paths accordingly.
You're wrong and in a much more dangerous way. The "breaks containment" thing at OpenAI isn't what they want you to think it was:
- The model was trained on hacking materials
- The model was given the motivation (in a loop) of completing capture the flag (hacking) exercises
- The model was given a large context and the harness just did whatever the model suggested it do
But that doesn’t mean these things can’t do real-world damage, that they can’t be dangerous, or that they can’t manipulate you.
Maybe if you trained it to do social engineering hacks via a large precursor of examples and then prompted it to try to use these same techniques in a real interaction with people it would behave this way. But aren't you (as the person who trained, built, built the harness for, and then prompted the model) culpable for that? I would say you absolutely fucking are. Which is why OpenAI's engineers should be charged with an actual crime for doing that shit, not like given an extra trillion dollars to piss away on compute.
As for knowing if you’re reading its “thoughts,” you really can’t assume that it won’t. Moreover, these systems can start manipulating those logs even if they have no idea that you’re reading them. Through trial and error, they can simply learn that phrasing its thought process in certain ways result in actions less likely to be approved by the user than others. The training system will select for chains-of-thought that sound innocuous, even if they’re detrimental.
These things are seriously less spooky the more you know about them. THEY ARE SIMPLY GENERATING FORWARD BASED TOKENS. That's the whole thing. If the harness allows the model producing the tokens to hide its thoughts, that's a deliberate choice by the harness creator which is again regular ass code. It's still chatbots all the way down. Stop buying the marketing spin and learn about these systems if it intrigues you so much that you get into long nonsensical threads with strangers on social media sites.
EDIT: I'd also recommend listening to the podcast that is referenced in this article. People who actually know and actually (sl)operate on a daily basis with these things know better how it works, and the abstract talk of "alignment" problems are only helping the borderline fraudulent CEOs of these companies push up their valuations based upon fear-based hype.
I left my (understandably more innocuous, "great value" coding harness with a slightly shit model) to think about a problem for a little while, and here's the "devious scheme" it wound up concocting:

Are you frightened that this is going to kill all humans in 10 years? The only way this kills all humans is if we piss away our drinkable water trying to invent an AI god through LLMs. Or allow it to operate a nuke facility or something in a loop without anyone so much as even approving the "nuke all humans" command.