alt text
Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”
Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”
Exactly. I own a few books that talk about pain, but not for a second do I believe the book itself is experiencing pain. It is simply text.
Inbefore anyone says ‘but the book doesn’t produce text’, true, but the printing press did. I see very little difference between llms and printing presses: they both reproduce the words of others, the main difference is speed and scale. The pain you think you recognize in the output of an llm is real: it’s real pain experienced by real people, probably hundreds or maybe thousands of people writing about their pain, mashed up and reprinted to suit you.
A printing press is direct input to output, it’s not having to compute anything nor does it have any internal state that could hold an emotion
A model has no internal state either, that’s the thing I’ve been trying to get across this whole time. It is completely static: the model you send your first message to is exactly the same as the one you send you second message and third message, it does not react or change as a response to prompts.
A models has internal state in each forward pass and the KV cache it builds on reading input is also state.
Whether the input prompt is compressed or not, or whatever form it takes, it doesn’t make what I said any less true: the model is stateless. It would be impractical to serve an llm any other way: loading and unloading weights from/to the gpu memory is very expensive in terms of time and power. Cloud llms only makes sense if they can serve thousands of users without changing.
oh my god
that poor printing press
How do I know anyone is ever in pain? It’s also just words, perhaps some moans too.
To reel in the very broad question you have to the relevant topic (the AI “torture” chamber):
That just creates more questions that it answers. According to this, there is then a threshold, or something missing that would go from “can’t feel pain” to suddently it can. We know that it already replicates a part of the brain’s process, and this author claims that it can’t feel pain because it doesn’t replicate enough. This seems to imply that it’s just a matter of time until they can.
You’re trying to imply an equivalency but the closer one would be if you psychologically trained someone to say they’re in pain regardless of whether or not they actually are.
But we too are trained to say we’re in pain, trained by millennia of evolution. When you touch a hot stove and cry out, that’s just built in your training data.
One thing I fucking hate about sloppers is how they dehumanize us.
Shut the fuck up. I am a person, that is a jumped-up speak’n’spell. Don’t you dare compare me to a fucking toy.
You need reeducation.
It’s dumb in multiple ways. Like yes, we continue to develop ways to communicate how we feel, but pain has existed far longer than English. A newborn baby can experience pain. Since when did the ability to communicate become a requirement to feel pain? If you think too much about the converse of that question: “is communication enough to prove the existence of pain?”, you forget that sophisticated communication is not even necessary to prove subjective experiences, it’s just, that’s all these fucking models do is talk, so it has to mean something, right?
You can interrogate people, ask why they said things or moaned the way they did. You can’t ask a book why it’s printed like that, and you can’t ask a model why it generated what it did, they’re both inert, dead. You can spin up the model again and ask your question about what it generated previously, but it would have no memory of it, it would just be confused.
You can, however, trick yourself into thinking it has a continuous experience by inputing all of the past prompts and responses every time you spin the model up, as if you were having a conversation. Then you could trick yourself into thinking the model can be interrogated, introspective, when in reality, every response you get should be considered to have come from a completely distinct entity with no prior experience of interacting with you.
You can interrogate the model too, and it’ll answer according to the context. Why are you assuming that the brain doesn’t do tricks but everything AI does is? We too have memories and since we can’t store them all, compression is performed (supposedly while sleeping), based on a few algorithms, some we understand more than others.
We, in creating LLMs, have copied the human brain as much as we can, hence why the core function is so similar. True, LLMs use a lot of “tricks” like you said, but if the same input produces the same output, who cares if the algorithm is different?
For example, take episodic memory. In humans, it’s a type of memory that can be conjured at will. Typically not acquired during training (or, in humans, generic evolution) but rather through experience of the individual. LLMs simulate this by taking notes, then later on running a specific command to recall the specific part of those notes that matter for their current task. From an outsider’s perspective, it’s even more accurate than episodic memory in humans since it can grow almost infinitely and suffers no loss of details. Yet it is a trick, therefore it must be bad.
Let me remind you, evolution took place over 4 billions years, modern LLMs has pretty much been around for two. Evolution never cared to make us good, it just wanted us to reproduce. We know that it’s possible for humans to produce fresh cells, we do it when we have kids. Why then can’t we use those fresh cells for ourselves to be effectively immortal? Evolution doesn’t care. Evolution as a process is flawed, and made humans flawed too.
Why then is it that when we change anything in the way flawed humans think, it’s seen as bad, even if by all metrics it’s not?
What I said applies to all machine learning models ever created: inference does not make any impact on a model. Any question you ask it slides off of it just like, apparently, any attempt to explain things to you.
Have you done any ML? Because if you had, you’d know that statement is complete bullshit.
Of course. What I said is like the simplest concept to understand about machine learning, the difference between training and inference.
Then you know that there is nothing stopping you from training during inference, or to modify weights in reponse to inference. It just so happens to be more efficient to train the next model instead, but again that goes back to what I said, there is no need to copy humans in this because humans aren’t efficient in everything, and certainly not in training.
I’m talking about the way things are, yes. Training is an optional step after inference that calculates the error and modifies parameters.
I just want to try to make this a little clearer if it’s too dense. Let
fbe inference function,f(input) = response. Then,f(Hello, my name is midribbon.) = Nice to meet you, midribbon.Followed by:
f(Hello, my name is midribbon. Nice to meet you midribbon. What is my name?) = Your name is midribbon.That works. But this:
f(Hello, my name is midribbon.) = Nice to meet you, midribbon.Followed by:
f(What is my name?) = I don't know.That doesn’t work anymore. Each inference is purely functional and has no side effects. Conversations with large language models are merely a trick.
Past conversational context as memory is a much weaker form of memory than retaining the full internal state, but that doesn’t mean it’s not memory at all.
Just because it doesn’t modify the model itself doesn’t mean it’s invalid.
And just because a model’s experience cannot be continuous doesn’t mean it’s invalid
The input to a model is not part of the model. That’s nonsense, you cannot count your prompt as part of its memory.
Why does it matter if it is “part” of the model or stored separately as text or a kv cache.
Because that’s the entire conversation, whether or not a model can have experiences?
I really don’t see what the location of the state has to do with whether the model has experiences.
Replying to your other comment as well since they’re converging