alt text
Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”
Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”
I just want to try to make this a little clearer if it’s too dense. Let
fbe inference function,f(input) = response. Then,f(Hello, my name is midribbon.) = Nice to meet you, midribbon.Followed by:
f(Hello, my name is midribbon. Nice to meet you midribbon. What is my name?) = Your name is midribbon.That works. But this:
f(Hello, my name is midribbon.) = Nice to meet you, midribbon.Followed by:
f(What is my name?) = I don't know.That doesn’t work anymore. Each inference is purely functional and has no side effects. Conversations with large language models are merely a trick.
Past conversational context as memory is a much weaker form of memory than retaining the full internal state, but that doesn’t mean it’s not memory at all.
Just because it doesn’t modify the model itself doesn’t mean it’s invalid.
And just because a model’s experience cannot be continuous doesn’t mean it’s invalid
The input to a model is not part of the model. That’s nonsense, you cannot count your prompt as part of its memory.
Why does it matter if it is “part” of the model or stored separately as text or a kv cache.
Because that’s the entire conversation, whether or not a model can have experiences?
I really don’t see what the location of the state has to do with whether the model has experiences.
Replying to your other comment as well since they’re converging
It’s not state though! Prompts are input, they belong to the user of the model. The model has absolutely no control over what gets sent to it.
By default, the previous conversation turns do. Just because it’s easier to corrupt than a human brain doesn’t make it not-state.
If ownership of the prompt is the issue, what about the model’s own output reasoning? What about notes/“memories” it may write for itself?
I feel like I’m talking to a wall. The entire conversation is whether the model can have experiences. The fact that the prompt comes from a user makes it external to the model. If you include the user as part of the system under investigation, of course you’ll come to the conclusion it can have experiences.