alt text
Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”
Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”
A printing press is direct input to output, it’s not having to compute anything nor does it have any internal state that could hold an emotion
A model has no internal state either, that’s the thing I’ve been trying to get across this whole time. It is completely static: the model you send your first message to is exactly the same as the one you send you second message and third message, it does not react or change as a response to prompts.
A models has internal state in each forward pass and the KV cache it builds on reading input is also state.
Whether the input prompt is compressed or not, or whatever form it takes, it doesn’t make what I said any less true: the model is stateless. It would be impractical to serve an llm any other way: loading and unloading weights from/to the gpu memory is very expensive in terms of time and power. Cloud llms only makes sense if they can serve thousands of users without changing.