alt text

Crudely drawn drawing of a person telling a computer “say ‘i am in pain’”. The computer replies with “> I AM IN PAIN”. The person then says “oh my god.”

yes, it’s a real thing

  • midribbon_action@lemmy.blahaj.zone
    link
    fedilink
    arrow-up
    2
    arrow-down
    1
    ·
    1 day ago

    You can interrogate people, ask why they said things or moaned the way they did. You can’t ask a book why it’s printed like that, and you can’t ask a model why it generated what it did, they’re both inert, dead. You can spin up the model again and ask your question about what it generated previously, but it would have no memory of it, it would just be confused.

    You can, however, trick yourself into thinking it has a continuous experience by inputing all of the past prompts and responses every time you spin the model up, as if you were having a conversation. Then you could trick yourself into thinking the model can be interrogated, introspective, when in reality, every response you get should be considered to have come from a completely distinct entity with no prior experience of interacting with you.

    • SorryQuick@lemmy.ca
      link
      fedilink
      arrow-up
      2
      ·
      1 day ago

      You can interrogate the model too, and it’ll answer according to the context. Why are you assuming that the brain doesn’t do tricks but everything AI does is? We too have memories and since we can’t store them all, compression is performed (supposedly while sleeping), based on a few algorithms, some we understand more than others.

      We, in creating LLMs, have copied the human brain as much as we can, hence why the core function is so similar. True, LLMs use a lot of “tricks” like you said, but if the same input produces the same output, who cares if the algorithm is different?

      For example, take episodic memory. In humans, it’s a type of memory that can be conjured at will. Typically not acquired during training (or, in humans, generic evolution) but rather through experience of the individual. LLMs simulate this by taking notes, then later on running a specific command to recall the specific part of those notes that matter for their current task. From an outsider’s perspective, it’s even more accurate than episodic memory in humans since it can grow almost infinitely and suffers no loss of details. Yet it is a trick, therefore it must be bad.

      Let me remind you, evolution took place over 4 billions years, modern LLMs has pretty much been around for two. Evolution never cared to make us good, it just wanted us to reproduce. We know that it’s possible for humans to produce fresh cells, we do it when we have kids. Why then can’t we use those fresh cells for ourselves to be effectively immortal? Evolution doesn’t care. Evolution as a process is flawed, and made humans flawed too.

      Why then is it that when we change anything in the way flawed humans think, it’s seen as bad, even if by all metrics it’s not?

      • midribbon_action@lemmy.blahaj.zone
        link
        fedilink
        arrow-up
        1
        ·
        1 day ago

        modern LLMs has pretty much been around for two [years]

        What I said applies to all machine learning models ever created: inference does not make any impact on a model. Any question you ask it slides off of it just like, apparently, any attempt to explain things to you.

        • SorryQuick@lemmy.ca
          link
          fedilink
          arrow-up
          2
          ·
          1 day ago

          Have you done any ML? Because if you had, you’d know that statement is complete bullshit.

            • SorryQuick@lemmy.ca
              link
              fedilink
              arrow-up
              2
              ·
              1 day ago

              Then you know that there is nothing stopping you from training during inference, or to modify weights in reponse to inference. It just so happens to be more efficient to train the next model instead, but again that goes back to what I said, there is no need to copy humans in this because humans aren’t efficient in everything, and certainly not in training.

                • SorryQuick@lemmy.ca
                  link
                  fedilink
                  arrow-up
                  1
                  ·
                  1 day ago

                  Does it matter whether it’s part of it or done immediately after? For all intents and purposes it’s the same thing. Like I said, if the input and outputs are the same, what does it matter how the process works?

                  From the user’s perspective, where one question results in many inference calls, it would look like the LLM learns while it works, assuming such training would be enabled, which they obviously wouldn’t but could do.

                  • midribbon_action@lemmy.blahaj.zone
                    link
                    fedilink
                    arrow-up
                    1
                    ·
                    1 day ago

                    I’m not talking about training, the paper this post is talking about is not about training, and no cloud llms allow their users to do training, so I have no idea what relevance it could have to this conversation. Maybe training is indeed really painful for llms, idk, that’s not what we’re talking about though.

    • midribbon_action@lemmy.blahaj.zone
      link
      fedilink
      arrow-up
      1
      ·
      1 day ago

      I just want to try to make this a little clearer if it’s too dense. Let f be inference function, f(input) = response. Then,

      f(Hello, my name is midribbon.) = Nice to meet you, midribbon.

      Followed by:

      f(Hello, my name is midribbon. Nice to meet you midribbon. What is my name?) = Your name is midribbon.

      That works. But this:

      f(Hello, my name is midribbon.) = Nice to meet you, midribbon.

      Followed by:

      f(What is my name?) = I don't know.

      That doesn’t work anymore. Each inference is purely functional and has no side effects. Conversations with large language models are merely a trick.

      • morrowind@lemmy.ml
        link
        fedilink
        arrow-up
        1
        arrow-down
        1
        ·
        1 day ago

        Past conversational context as memory is a much weaker form of memory than retaining the full internal state, but that doesn’t mean it’s not memory at all.

        Just because it doesn’t modify the model itself doesn’t mean it’s invalid.

        And just because a model’s experience cannot be continuous doesn’t mean it’s invalid

          • morrowind@lemmy.ml
            link
            fedilink
            arrow-up
            1
            ·
            1 day ago

            Why does it matter if it is “part” of the model or stored separately as text or a kv cache.

              • morrowind@lemmy.ml
                link
                fedilink
                arrow-up
                1
                ·
                1 day ago

                I really don’t see what the location of the state has to do with whether the model has experiences.

                Replying to your other comment as well since they’re converging

                  • morrowind@lemmy.ml
                    link
                    fedilink
                    arrow-up
                    1
                    ·
                    24 hours ago

                    By default, the previous conversation turns do. Just because it’s easier to corrupt than a human brain doesn’t make it not-state.

                    If ownership of the prompt is the issue, what about the model’s own output reasoning? What about notes/“memories” it may write for itself?