Musk denied he was aware Grok ever produced ‘any naked underage images’

A survivor of child sexual abuse has sued Elon Musk’s artificial intelligence company, alleging that its chatbot used pictures of her abuse to generate new illegal pornographic images that depict her.

“Using real images of Plaintiff and class members, Grok generated child pornography depicting Plaintiff and class members,” states the complaint, which was filed last week in a US district court in California.

Attorneys for the plaintiff, who is listed as Jane Doe in the case to protect her identity, accuse xAI of both generating CSAM of their client and of ingesting child sexual abuse images depicting her into the company’s datasets after new images were publicly posted.

  • Digestive_Biscuit@feddit.uk
    link
    fedilink
    English
    arrow-up
    5
    ·
    1 day ago

    How does their AI engines (whatever they are called) access such material? I’m assuming its from the dark web or a government library or something and if that’s the case then somebody made a decision to include such questionable sources.

    • partofthevoice@lemmy.zip
      link
      fedilink
      arrow-up
      4
      ·
      18 hours ago

      It’s more like if you kept the recipe to the photo, rather than the photo itself. The AI model, when trained, is having its “weights” updated — which means they’re tuning a very long set of lists of numbers.

      For example, imagine:

      (
        [0.028474, 0.274729, …],
        [0.827482, 0.283759, …],
        …*billions of lists
      )
      

      The numbers are actually random generated at first. Training can involve tricks like cutting out pieces of an image, then telling an AI to predict what goes in the empty space. Or reducing a photos quality, then telling the AI to increase its quality. The important part is that you distort the image while retaining the original as the answer key.

      For each answer, you measure how correct/incorrect the model was. If it’s correct, you don’t do anything. If it’s incorrect, you do mathematical tricks to update those numbers.

      Those numbers bias the model, all the way from random gibberish output to coherent output. So they go through the process I mentioned before billions of times, each time with different pictures / distortions, each time ever so slightly nudging those numbers until the model spits out coherent outputs.

      Those numbers wind up being something like a recipe. Like if you’d stored the exact pixel configuration to a photo, but didn’t actually have the photo. Except this recipe tries to measure the general semantic relationships between words and images, such that it can generate images when given prompts. This recipe is less deterministic, but nonetheless it’s tainted with CP shit.

    • Ithral@lemmy.blahaj.zone
      link
      fedilink
      arrow-up
      6
      ·
      1 day ago

      Model weights is what they are called. They didnt access them exactly, if the allegations are true, what hapoened was during training those images were downloaded and used to train the image generation weights. So the problem here is the training data was not carefully sanitized and was just randomly downloaded from everywhere resulting in illegal content slipping in.

  • zecg@lemmy.world
    link
    fedilink
    arrow-up
    26
    ·
    2 days ago

    The distinction of using pre-existing CSAM is notable because law enforcement and child protection organizations frequently give those illegal materials a kind of digital fingerprint – known as a hash – which allows them to track images and videos when they appear online. In this case, attorneys for the plaintiff stated that the Canadian Centre for Child Protection used images’ fingerprints to identify AI-generated CSAM on X that depicted their client.

    How does that work? Are there hashes of child sex abuse images somehow reproduced by LLM? Surely the new images don’t have the same hashes. Guardian is not really being clear here.

    edit: Ars’ is less badly written: “This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.””

    • cantstopthesignal@sh.itjust.works
      link
      fedilink
      arrow-up
      9
      arrow-down
      2
      ·
      2 days ago

      From my understanding a hash completely changes when a single part of it is changed. That’s the point of a hash. Sounds like they identified the hashed images from the training set.

      • Clent@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        18
        ·
        2 days ago

        Nope. They use perceptual hashing. That way water marking and cropped images still trigger against the casm scanners.

        The known casm results in a specific hash. Modified versions of that will result in a different hash but such that the “distance” is meaningful. The distance is how many bits need to flip to match the known casm hash.

        It can lead to false positives so that’s why manually checking occurs. It’s also why people got very upset that Apple and perhaps others were going to automatically do these checks during upload to the cloud. When performed against billions of images, the false positive rate becomes huge and manually checking becomes too much overhead so people would be falsely accused.

        It’s very likely what they found in groks models is a false positive but in this case they can probably compel twitter to prove it’s a false positive.

        • Cypher@aussie.zone
          link
          fedilink
          arrow-up
          2
          ·
          21 hours ago

          It’s also why people got very upset that Apple and perhaps others were going to automatically do these checks during upload to the cloud.

          They were planning on doing checks locally, potentially impacting device performance, battery life and privacy as false positives would still be reviewed by humans. Potentially leading to individuals intimate photos being viewed without consent.

      • FiskFisk33@startrek.website
        link
        fedilink
        arrow-up
        1
        ·
        1 day ago

        there are other, more robust, ways of fingerprinting images than hashing the data. I’m not sure how it works exactly though, but I can imagine it is not completely different from how shazam can pick out a song in a noisy environment.

    • Grimy@lemmy.world
      link
      fedilink
      arrow-up
      3
      ·
      2 days ago

      Ya it doesn’t make sense. I think it’s probably facial recognition and since she’s in the training dataset, her face pops up when the model gets pulled in that direction.

      They wouldn’t have access to the dataset, I don’t see how they could know for the Arts explanation. Maybe it learned the hash and reproduces it but that doesn’t seem likely. Even the part about adding hashs to track the images doesn’t make much sense. That would imply they are distributing it somehow.

      I don’t get why X isn’t running age detection on the output though.

  • RagingRobot@lemmy.world
    link
    fedilink
    arrow-up
    17
    ·
    2 days ago

    If he was unaware why is he sewing to block a law restricting this? How could he possibly be unaware of what’s going on inside his own and the whole world knows?