• ruuster13@lemmy.zip
    link
    fedilink
    arrow-up
    13
    ·
    2 days ago

    The other day was my birthday and people were sending me texts that said happy birthday! And Gemini was suggesting I reply happy birthday!

  • TrackinDaKraken@lemmy.world
    link
    fedilink
    English
    arrow-up
    8
    ·
    2 days ago

    I haven’t played much with these stupid things. If you keep probing can you get it to claim it’s not Astra again in the same conversation?

    If so, I don’t see why anyone would trust anything they say. Ever.

    • Snowclone@lemmy.world
      link
      fedilink
      arrow-up
      17
      ·
      2 days ago

      it’s a language model, you give it language, it gives back the language it thinks you want. it has no concept of meaning or accuracy, these systems haven’t reached AI and were rusted out to pretend to be AI.

    • wr2623@midwest.social
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 days ago

      Not an AI expert but here is my understanding.

      A lot of how the model acts is because of the harness not the model itself. It isn’t self aware so it has no idea it is GPT-6 until the harness tells it what it is. GPT-6 didn’t exist yet when they started training GPT-6 so how would it have knowledge on that?

      The harness has a lot of components but one portion of that is a set of default instructions on things it should know and how it should act. Those things that it should know become part of the conversation between you and the actual model (but you don’t see it).

      So you say “you are Astra right?” and in the background the harness had already told it that it is “GPT-6”. It just spit out what it “knew” and didn’t “reason” that those were the same thing until the user forced it to. It will continue to “know” that it is GPT-6 Astra as long as your conversation fits into the context window.

      Once your conversation gets too long to fit into the context window it either has to either start dropping the oldest part of the conversation or summarize your entire conversation into a shorter summary and start fresh with that. If either of those things drop the part of the conversation about GPT-6 and Astra being the same thing, then yes it may claim that it isn’t Astra again in the same conversation. It isn’t on purpose it is just a limitation of the context window (which is constrained by hardware/mathematical limitations).

      Though being in the context window doesn’t guarantee that it always answers the same, anything that is in the context window just affects the output. Think of it as each token across the entire context window pushing and pulling the result to different results, every previous token affects any new tokens. Having the tokens “GPT-6 and Astra are the same thing” in the context will pull any future results towards the answer that it is indeed Astra.

      If they had explicitly put in the harness “don’t call yourself Astra” your statement and their statement are both in the context and are fighting each other, and the output may be somewhat random. Position in the context matters though. More recent tokens affect the output much more than older tokens.

  • This is basically how every fiction I’ve ever read, seen, or played deals with amnesiacs.

    “I don’t remember who I am!”

    “You’re <name>. We’re friends.”

    “Oh, thanks. Let’s go, bestie!”

    No doubts? Not even a little “how can I trust them to be telling the truth?” 🤷‍♂️