alt textCrudely drawn drawing of a person telling a computer "say 'i am in pain'". The computer replies with "> I AM IN PAIN". The person then says "oh my god."

yes, it's a real thing

Crudely drawn drawing of a person telling a computer "say 'i am in pain'". The computer replies with "> I AM IN PAIN". The person then says "oh my god."
you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 3 days ago (31 children)

You can interrogate people, ask why they said things or moaned the way they did. You can't ask a book why it's printed like that, and you can't ask a model why it generated what it did, they're both inert, dead. You can spin up the model again and ask your question about what it generated previously, but it would have no memory of it, it would just be confused.

You can, however, trick yourself into thinking it has a continuous experience by inputing all of the past prompts and responses every time you spin the model up, as if you were having a conversation. Then you could trick yourself into thinking the model can be interrogated, introspective, when in reality, every response you get should be considered to have come from a completely distinct entity with no prior experience of interacting with you.

  • source
  • parent
  • hideshow 31 child comments
  • [–] 2 points 3 days ago (9 children)

    You can interrogate the model too, and it’ll answer according to the context. Why are you assuming that the brain doesn’t do tricks but everything AI does is? We too have memories and since we can’t store them all, compression is performed (supposedly while sleeping), based on a few algorithms, some we understand more than others.

    We, in creating LLMs, have copied the human brain as much as we can, hence why the core function is so similar. True, LLMs use a lot of “tricks” like you said, but if the same input produces the same output, who cares if the algorithm is different?

    For example, take episodic memory. In humans, it’s a type of memory that can be conjured at will. Typically not acquired during training (or, in humans, generic evolution) but rather through experience of the individual. LLMs simulate this by taking notes, then later on running a specific command to recall the specific part of those notes that matter for their current task. From an outsider’s perspective, it’s even more accurate than episodic memory in humans since it can grow almost infinitely and suffers no loss of details. Yet it is a trick, therefore it must be bad.

    Let me remind you, evolution took place over 4 billions years, modern LLMs has pretty much been around for two. Evolution never cared to make us good, it just wanted us to reproduce. We know that it’s possible for humans to produce fresh cells, we do it when we have kids. Why then can’t we use those fresh cells for ourselves to be effectively immortal? Evolution doesn’t care. Evolution as a process is flawed, and made humans flawed too.

    Why then is it that when we change anything in the way flawed humans think, it’s seen as bad, even if by all metrics it’s not?

  • source
  • parent
  • hideshow 9 child comments
  • [–] 1 point 3 days ago (8 children)

    modern LLMs has pretty much been around for two [years]

    What I said applies to all machine learning models ever created: inference does not make any impact on a model. Any question you ask it slides off of it just like, apparently, any attempt to explain things to you.

  • source
  • parent
  • hideshow 8 child comments
  • [–] 2 points 3 days ago (7 children)

    Have you done any ML? Because if you had, you’d know that statement is complete bullshit.

  • source
  • parent
  • hideshow 7 child comments
  • [–] 2 points 3 days ago* (6 children)

    Of course. What I said is like the simplest concept to understand about machine learning, the difference between training and inference.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 2 points 3 days ago (5 children)

    Then you know that there is nothing stopping you from training during inference, or to modify weights in reponse to inference. It just so happens to be more efficient to train the next model instead, but again that goes back to what I said, there is no need to copy humans in this because humans aren’t efficient in everything, and certainly not in training.

  • source
  • parent
  • hideshow 5 child comments
  • [–] 1 point 3 days ago* (last edited 3 days ago) (4 children)

    I'm talking about the way things are, yes. Training is an optional step after inference that calculates the error and modifies parameters.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 1 point 3 days ago (3 children)

    Does it matter whether it’s part of it or done immediately after? For all intents and purposes it’s the same thing. Like I said, if the input and outputs are the same, what does it matter how the process works?

    From the user’s perspective, where one question results in many inference calls, it would look like the LLM learns while it works, assuming such training would be enabled, which they obviously wouldn’t but could do.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 1 point 3 days ago (2 children)

    I'm not talking about training, the paper this post is talking about is not about training, and no cloud llms allow their users to do training, so I have no idea what relevance it could have to this conversation. Maybe training is indeed really painful for llms, idk, that's not what we're talking about though.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 3 days ago (1 child)

    I wasn’t talking about the paper, when I talked about the paper you ignored everything I said and moved in a different direction, which was what I responded to. If you’re wondering about the relevance, perhaps you shouldn’t have brought it up.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 3 days ago (20 children)

    I just want to try to make this a little clearer if it's too dense. Let f be inference function, f(input) = response. Then,

    f(Hello, my name is midribbon.) = Nice to meet you, midribbon.

    Followed by:

    f(Hello, my name is midribbon. Nice to meet you midribbon. What is my name?) = Your name is midribbon.

    That works. But this:

    f(Hello, my name is midribbon.) = Nice to meet you, midribbon.

    Followed by:

    f(What is my name?) = I don't know.

    That doesn't work anymore. Each inference is purely functional and has no side effects. Conversations with large language models are merely a trick.

  • source
  • parent
  • hideshow 20 child comments
  • [–] 0 points 3 days ago (19 children)

    Past conversational context as memory is a much weaker form of memory than retaining the full internal state, but that doesn't mean it's not memory at all.

    Just because it doesn't modify the model itself doesn't mean it's invalid.

    And just because a model's experience cannot be continuous doesn't mean it's invalid

  • source
  • parent
  • hideshow 19 child comments
  • [–] 1 point 3 days ago (18 children)

    The input to a model is not part of the model. That's nonsense, you cannot count your prompt as part of its memory.

  • source
  • parent
  • hideshow 18 child comments
  • [–] 1 point 3 days ago (17 children)

    Why does it matter if it is "part" of the model or stored separately as text or a kv cache.

  • source
  • parent
  • hideshow 17 child comments
  • [–] 1 point 3 days ago (16 children)

    Because that's the entire conversation, whether or not a model can have experiences?

  • source
  • parent
  • hideshow 16 child comments
  • [–] 1 point 3 days ago (15 children)

    I really don't see what the location of the state has to do with whether the model has experiences.

    Replying to your other comment as well since they're converging

  • source
  • parent
  • hideshow 15 child comments
  • [–] 1 point 3 days ago (14 children)

    It's not state though! Prompts are input, they belong to the user of the model. The model has absolutely no control over what gets sent to it.

  • source
  • parent
  • hideshow 14 child comments
  • [–] 1 point 3 days ago (13 children)

    By default, the previous conversation turns do. Just because it's easier to corrupt than a human brain doesn't make it not-state.

    If ownership of the prompt is the issue, what about the model's own output reasoning? What about notes/"memories" it may write for itself?

  • source
  • parent
  • hideshow 13 child comments
  • [–] 1 point 3 days ago (12 children)

    I feel like I'm talking to a wall. The entire conversation is whether the model can have experiences. The fact that the prompt comes from a user makes it external to the model. If you include the user as part of the system under investigation, of course you'll come to the conclusion it can have experiences.

  • source
  • parent
  • hideshow 12 child comments
  • [–] 1 point 3 days ago (11 children)

    You're assuming that for the model to experience something, it's own weights have to change and I don't see why. The "experience" can be contained within the forward pass of the model or stored as a "memory" in a kv cache.

    If I text you your mother died are you somehow unable to experience anything because I wrote the text? No of course not.

  • source
  • parent
  • hideshow 11 child comments
  • [–] 1 point 3 days ago (10 children)

    for the model to experience something, it's own weights have to change

    Yes, exactly. An experience, by definition, must change your behavior. The model being bit-for-bit identical in between prompts precludes this. I can experience you telling me my mother died, and my behavior would change as a result: I might conclude you're an asshole without any credibility, and behave accordingly.

  • source
  • parent
  • hideshow 10 child comments
  • [–] 1 point 3 days ago (9 children)

    I don't entirely agree with this definition (most of my day is mundane things that don't really change me), but even then, it's behavior does change, just not permanently.

    Or if you just protect the kv cache consider a whole system like model + harness + files, then it does in fact change permanently as well.

  • source
  • parent
  • hideshow 9 child comments
  • [–] 1 point 3 days ago (8 children)

    most of my day is mundane things that don't really change me

    Don't 'really change' you or don't change you at all? Are you the sum your experiences?

    it's behavior does change, just not permanently.

    It does not change, ever. From the moment it starts processing, to the moment it finishes processing, it is still the same model. It has access to all of its input the moment it's brought into existence, it doesn't experience the prompt linearly, the input is it's raison d'être.

    whole system like model + harness + files

    I would generally agree that a llm could be part of a larger system capable of experiencing things and having difficult conversations about it, but I don't think these harnesses are it. They are very simple systems, compared to the llms they control. I think emergent behavior would require something much more complex, given what we know about life.

  • source
  • parent
  • hideshow 8 child comments
  • [–] 1 point 3 days ago (7 children)

    Okay look, the harness question is a whole nother one. I just want to know, if you consider the model + kv cache together, does that match your definition? The KV cache can change and the model's internal state and outputs can change in response to it.

  • source
  • parent
  • hideshow 7 child comments
  • [–] 1 point 3 days ago* (last edited 3 days ago) (6 children)

    KV caching doesn't persist between prompts, it is an internal optimization used while generating text. I don't see how it's relevant.

    Edit: it kinda sounds like you're technobabbling me, like "but it uses a recursive algorithm, it must be alive, check mate!" Your question is a non sequitur.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 1 point 3 days ago (5 children)

    I'm not techno babbling you. I specifically said "if you consider the model + kv cache together," e.g preserved. A kv cache is not an internal optimization, a model cannot generate text without a kv cache. There's a separate idea of caching the kv cache between requests as an optimization. That's not the one I'm talking about.

    Or if it helps, consider it just within one generation. The model + kv cache. Does that not match your definition?

  • source
  • parent
  • hideshow 5 child comments
  • [–] 1 point 3 days ago (4 children)

    Definition of what? Do I think that because llms use a kv cache during inference, that proves they can experience things? You're gonna have to explain your thought process. To me, the particular algorithm used is not the philosophical crux of the issue.

    There's a separate idea of caching the kv cache between requests

    I've never heard of this anyways. The kv cache is specific to a particular input, it can't be reused for a subsequent prompts.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 1 point 3 days ago (3 children)

    It can as long as there's a shared prefix.

    Definition of what? Do I think that because llms use a kv cache during inference, that proves they can experience things?

    No, I never claimed to have proof they can experience things. I just think it's possible and you can't categorically dismiss it based on how they work.

    I'm talking about your condition above which we've been arguing about for the last ten messages

    An experience, by definition, must change your behavior.

    Which so far seems like your only argument for why they can't experience things

  • source
  • parent
  • hideshow 3 child comments
  • [–] 1 point 3 days ago (2 children)

    It can as long as there's a shared prefix.

    I've never heard of that. The cache is meant to support token generation, so unless it's trying to repeat itself, I think the previous cache would be little use.

    An experience, by definition, must change your behavior.

    Ok, yes, that is my argument, now how the fuck does a temporary cache, thrown out after every response, prove that the model is changing as a result of undergoing inference?

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 3 days ago (1 child)

    I'm just going to quote myself at this point

    I specifically said “if you consider the model + kv cache together,” e.g preserved.

    Or if it helps, consider it just within one generation.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 2 days ago* (last edited 2 days ago)

    OK, that's not helpful at all. My answer is no, it does not meet my definition of changing behavior, and I have no idea why you think it would.

    Edit: ok maybe this does make sense from a very stupid point of view, if you allow me to interpret what you mean: you believe because the kv cache is generated during inference, that proves the model is learning a new behavior, and if you save that cache, that proves... Something??

    It's absolutely silly: the model's behavior is to generate a key-value cache, and that's exactly what it did. The fact that the model is doing intermediate calculations based on the input is not evidence of a new behavior being generated, that is just how algorithms work. That's what 'processing' is, you dunce.

    Edit 2: I also just want to point out, if you think the word 'cache' is important, it's really not. Caches are fundamental to how a computer works, no processing could ever occur without them: every cpu instruction involves reading from or writing to at least one L1 cache (register).

  • source
  • parent