Does anyone find it at all mechanistically surprising that one of the many latent dimensions of the set of all text humans produce is 'subject is in pain' or that pushing activations in that direction will alter the output of a text prediction system?
Reminds me of the way that when you relax the finetuning on LLMs they start 'believing' in UFOs and supernatural stuff. And how people take that to somehow mean something. Of course that is a large component in the space of stuff humans write about that you have to suppress if you want a professionally useful text service!
The research is cool. I love all the work that has been done decomposing representations within these systems.
Edit: Damn. I am forced to agree with Mustafa Suleyman of Microsoft.
"AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans,” Suleyman wrote. “Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings […] If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity."
"AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do."