I'm not suggesting it will try and bypass the stop button. I’m saying it will try and manipulate you into approving something you wouldn’t otherwise. And yes, you can read some of its “thoughts,” but it knows you’re reading them, or it can determine that through trial and error. And then it can start subtly manipulating those recorded “thoughts” to make them sound different from what they really are.
-
No "it" doesn't know that you're reading "its" "thoughts".
-
No "it" wouldn't, because "it" is just generating plausible text and has no motivations of "its" own.
-
There is no "it".
LLMs are still functionally useless without a harness and do nothing useful without tools like an MCP server.
The "thought" bubbles are no different from the rest of the plausible text "it" is generating. They're so indistinguishable to "it" that that's an attack surface for injection attacks.
EDIT: Many of the biggest forward breakthroughs in LLM coding (or vibe coding) have come from harness improvements, not model improvements. In many harnesses, the models themselves are able to be substituted mid-session. Models work better in my experience when a human actively steers them away from stupid ideas by reading their thoughts and occasionally interrupting them. I have a few slopjects that I'm sloperating upon right now, and I can get results out of "great value" claude code (opencode) using this approach, even if it sometimes goes completely "off the rails" and does shit like saying "retained" over and over again until the harness pulls the plug.