I'm just going to quote myself at this point
I specifically said “if you consider the model + kv cache together,” e.g preserved.
Or if it helps, consider it just within one generation.
I'm just going to quote myself at this point
I specifically said “if you consider the model + kv cache together,” e.g preserved.
Or if it helps, consider it just within one generation.
It can as long as there's a shared prefix.
Definition of what? Do I think that because llms use a kv cache during inference, that proves they can experience things?
No, I never claimed to have proof they can experience things. I just think it's possible and you can't categorically dismiss it based on how they work.
I'm talking about your condition above which we've been arguing about for the last ten messages
An experience, by definition, must change your behavior.
Which so far seems like your only argument for why they can't experience things
I'm not techno babbling you. I specifically said "if you consider the model + kv cache together," e.g preserved. A kv cache is not an internal optimization, a model cannot generate text without a kv cache. There's a separate idea of caching the kv cache between requests as an optimization. That's not the one I'm talking about.
Or if it helps, consider it just within one generation. The model + kv cache. Does that not match your definition?
Okay look, the harness question is a whole nother one. I just want to know, if you consider the model + kv cache together, does that match your definition? The KV cache can change and the model's internal state and outputs can change in response to it.
I don't entirely agree with this definition (most of my day is mundane things that don't really change me), but even then, it's behavior does change, just not permanently.
Or if you just protect the kv cache consider a whole system like model + harness + files, then it does in fact change permanently as well.
You're assuming that for the model to experience something, it's own weights have to change and I don't see why. The "experience" can be contained within the forward pass of the model or stored as a "memory" in a kv cache.
If I text you your mother died are you somehow unable to experience anything because I wrote the text? No of course not.
By default, the previous conversation turns do. Just because it's easier to corrupt than a human brain doesn't make it not-state.
If ownership of the prompt is the issue, what about the model's own output reasoning? What about notes/"memories" it may write for itself?
I really don't see what the location of the state has to do with whether the model has experiences.
Replying to your other comment as well since they're converging
Why does it matter if it is "part" of the model or stored separately as text or a kv cache.
A models has internal state in each forward pass and the KV cache it builds on reading input is also state.
I've already seen advertising agencies use llms as simulated customers. No idea how well it works
Past conversational context as memory is a much weaker form of memory than retaining the full internal state, but that doesn't mean it's not memory at all.
Just because it doesn't modify the model itself doesn't mean it's invalid.
And just because a model's experience cannot be continuous doesn't mean it's invalid