alt textCrudely drawn drawing of a person telling a computer "say 'i am in pain'". The computer replies with "> I AM IN PAIN". The person then says "oh my god."

yes, it's a real thing

Crudely drawn drawing of a person telling a computer "say 'i am in pain'". The computer replies with "> I AM IN PAIN". The person then says "oh my god."
you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 3 days ago (2 children)

A model has no internal state either, that's the thing I've been trying to get across this whole time. It is completely static: the model you send your first message to is exactly the same as the one you send you second message and third message, it does not react or change as a response to prompts.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 3 days ago (1 child)

    A models has internal state in each forward pass and the KV cache it builds on reading input is also state.

  • source
  • parent
  • hideshow 1 child comment
  • Whether the input prompt is compressed or not, or whatever form it takes, it doesn't make what I said any less true: the model is stateless. It would be impractical to serve an llm any other way: loading and unloading weights from/to the gpu memory is very expensive in terms of time and power. Cloud llms only makes sense if they can serve thousands of users without changing.

  • source
  • parent