• gwheel@lemmy.zip
    link
    fedilink
    English
    arrow-up
    72
    ·
    24 days ago

    Assuming a checker tool is public it’s enough for a teacher to verify classwork, but AI providers having the sole ability to identify generated content with no way to independently verify is not a real solution.

    Plus this site advertises a tool to remove this watermarking, so it can’t be that hard to scrub out if you’re aware of it.

    • Leon@pawb.social
      link
      fedilink
      English
      arrow-up
      57
      arrow-down
      1
      ·
      24 days ago

      The goal is to ensure that they don’t inbreed their models, not fix the problems they’ve caused.

      • brucethemoose@lemmy.world
        link
        fedilink
        English
        arrow-up
        24
        ·
        24 days ago

        This won’t fix the inbreeding issue, anyway. The bias is extremely slight, but random, and orthogonal to Claude’s own “slop patterns” and tendencies. And theres tons of other LLM content that will end up in their dataset outside their control.

        Besides, as much as Claude accusess others of it, everyone’s training on everyone else’s output and they know it.

    • Diurnambule@jlai.lu
      link
      fedilink
      English
      arrow-up
      7
      ·
      24 days ago

      I wonder what would happen if some start to watermark document they doesn’t want in Claude training

      • turtlesareneat@piefed.ca
        link
        fedilink
        English
        arrow-up
        5
        ·
        24 days ago

        There’s a key involved that we don’t have, so people can’t do this on their own. It’s pretty fascinating. Training models would have to be told to check for watermarks and ignore them, but yeah that would be an effective way for the providers to avoid ingesting their own AI output.

  • Opisek@piefed.blahaj.zone
    link
    fedilink
    English
    arrow-up
    55
    arrow-down
    1
    ·
    24 days ago

    I thought declaude would be like degoogle.

    Nope, turns out what they do is precisely the opposite. They try to remove those described markers. Real scummy.

      • 🌞 Alexander Daychilde 🌞@lemmy.world
        link
        fedilink
        English
        arrow-up
        13
        arrow-down
        2
        ·
        24 days ago

        I think the thing that frustrates me most about AI is that so many people seem to have forgotten that writers exist and wrote huge amounts of text - books, news articles, magazines - for centuries. But now, anytime someone writes a few cogent paragraphs, you get people screaming “AI!!! IT’S AI!!!”.

        I know, because it has happened to me.

        The article, to me, sounds like a competent author wrote it. Which is representative of a lot of the text LLMs were trained on.

        • black0ut@pawb.social
          link
          fedilink
          English
          arrow-up
          10
          ·
          24 days ago

          I can’t tell you about the text, but the whole site is vibe coded. The CSS, the js, everything. Which doesn’t really inspire much confidence on the text, especially as it’s explaining how to bypass the claude detector.

    • brucethemoose@lemmy.world
      link
      fedilink
      English
      arrow-up
      11
      arrow-down
      3
      ·
      24 days ago

      I’m against trying to “censor” LLMs, but yeah. What’s even the ostensible benevolent scenario for stripping invisible watermarking from Claude?

      I can’t think of one, playing devil’s advocate.

      It’s pretty scummy, indeed.

      • AlmightyDoorman@kbin.earth
        link
        fedilink
        arrow-up
        7
        ·
        24 days ago

        I am against AI and LLMs but the reason is quite simple, a lot of modern writing is offical bullshit to get stuff approved. Research grants, medical therapy approval, offical work mail that has to sound professional, job applications (insofar that noone cares what you write, they want a standard template and pick a candidate along qualifications anyways.). All these texts cost a lot of time and it makes no difference if a human or some copy machine writes it. And tbh. were it not for all the disadvantages of modern LLMs i would use it for the exact same reason, because sitting at a grant application for two hours while others do it in 5min is fucking useless and exhausting.

        • CovertOperative@piefed.zip
          link
          fedilink
          English
          arrow-up
          9
          ·
          24 days ago

          If it’s stuff that no one cares what you write, then they also won’t care if it’s proveably written by AI or not. So removing the watermark still does nothing beneficial.

        • brucethemoose@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          2
          ·
          edit-2
          24 days ago

          Okay.

          I don’t agree. But let’s say I agree.

          …Just don’t use Claude?

          Use an LLM without a watermark; there are hundreds to pick from.


          In other words, if one is going to try to hide automated writing, I think there should be a bare minimum effort to do so. That includes:

          • Reading/checking the text, to see if it makes any sense.

          • Actually trying to pass it as human.

          90% of slop is brain melting slop because this minimum bar isn’t even met. And all Claude’s watermark would do is catch that bottom of the barrel; it wouldn’t censor anyone.

          • tom@jlai.lu
            link
            fedilink
            English
            arrow-up
            2
            ·
            24 days ago

            From what I understood you need statistical tests to detect the watermark. You cannot easily detect by eyeballing that a coin will land on heads 55% of the time instead of 50%

  • Optional@lemmy.world
    link
    fedilink
    English
    arrow-up
    45
    arrow-down
    2
    ·
    24 days ago

    Well shit, that was fuckin’ interesting.

    Yes, AI is evil. So is facebook. But the engineering is still interesting.

    • Thorry@feddit.org
      link
      fedilink
      English
      arrow-up
      31
      ·
      24 days ago

      That’s one of the things that frustrates me most about this whole AI thing. I fucking hate it and I want it to die, I wish it were never created in the first place. But from a tech enthusiast and a maths nerd point of view, it is super interesting.

      Like the performance of these models is shit compared to a real person doing actual work. But if we think about what we are doing on a basic level, the performance is way beyond what I would expect it to be. I wouldn’t expect it to be able to form a coherent sentence or scale as well as it does (even though the resources required to run these is still very high).

      It could have been really cool shit people did studies on and played around with to explore the math. Cool little play models we could let go on a bunch of data and see what it did and how. Something for a small group of nerds and experts who are into that kind of thing, for the sake of learning and nothing else.

      But no, somehow it got transmorphed into “AI”. And marketed like this actual learning almost sentient computer system that can replace all workers. You can ask it anything and it will give PhD level expert answers. Oh and it’s run by a handful of the most vile men imaginable who pour all of the world’s money and resources into it, all so they get to be god emperor of the world. Fucking terrible.

      • frongt@lemmy.zip
        link
        fedilink
        English
        arrow-up
        6
        arrow-down
        1
        ·
        24 days ago

        Yeah, machine learning is incomprehensibly impressive! But it’s the implementation where corporations have slurped up everyone’s work and turned it into private profit while also wrecking every kind of media that fucking sucks.

        Like if a stock photo company wanted to train and use a model for describing stuff in their library, great! Tagging, describing, and enabling discovery is a difficult task. But using it to slop out some low-quality images? You should reconsider what you’re doing with your life.

      • melfie@lemmy.zip
        link
        fedilink
        English
        arrow-up
        3
        ·
        23 days ago

        The emergent behavior in huge models where it can “reason” instead of simply predicting the next token is fascinating. Artificial “neurons” built on statistics and linear algebra emerge to create something that legitimately has artificial intelligence. ANNs were conceived in the 1940s, building on centuries of development in statistical modeling and only now do we have the compute power to make this vision a reality.

        Yes, the “intelligence” has significant limitations and won’t be replacing human intelligence anytime soon, but it can actually be a useful tool if its limitations are kept in mind.

        The problem is of course the tech bros turning centuries of innovation they had no part in developing into a massive Ponzi scheme for their own profit.

      • Optional@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        1
        ·
        24 days ago

        Yeah totally agree. And I guess it’s all down to the astronomical amount of guesses it gets to make in a given second. Sort of like contemplating infinity, but with words and testable.

    • Scrubbles@poptalk.scrubbles.tech
      link
      fedilink
      English
      arrow-up
      9
      ·
      24 days ago

      AI isn’t evil. Generative AI isn’t evil. AI has existed for 20+ years now, I studied it back in my uni days.

      Corporations, how they trained it, how they use it now, how they are willing to pave the planet to force it down our throats is evil.

      This is one of those things as tech people we have to come to terms with and understand. No technology is inherently good or evil, it’s what people do with it.

  • MagicShel@lemmy.zip
    link
    fedilink
    English
    arrow-up
    14
    arrow-down
    1
    ·
    24 days ago

    I’m not convinced that they even know 100% how Anthropic is doing it. I can think of an easier way that doesn’t corrupt the text: just find a bunch of tokens where there is a good spread of token possibilities, and the more often the most likely one is chosen, the more likely it’s AI.

    That being said, it doesn’t seem much different from what any of us do to identify AI text — it has lots of tells anyway.

    • Hildegard (she/her)@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      13
      arrow-down
      1
      ·
      24 days ago

      That’s how AI testers work and its why they don’t. Most forms of formal writing are predictable by design. If the AI can predict predictable formulaic writing, it doesn’t mean its AI, its probably just any form of professional writing other than fiction.

      Famous public domain works will always be considered AI by those tests, because of course your LLM knows the american national constitution. It was in the training data, so it can predict it with 100% accuracy, therefore your test wrongly calls it AI.

      Testing for AI writing that way does not work.

      • MagicShel@lemmy.zip
        link
        fedilink
        English
        arrow-up
        4
        ·
        24 days ago

        The difference between what you describe and what I describe, is that a 100% match isn’t a hit. Nor is a 90/7/2/1. You need something with meaningful variability. Even within formal papers there are places where word choice is arbitrary as the article explains.

        Of course, you’re lacking the context of the full prompt and just feeding in the raw text. Again it gets way more reliable the more text you have.

        But it’s moot because the more text you have the more tells will sneak in and you probably don’t even need an AI checker. Those phrases that AI loves but humans use comparatively rarely. It’s not a tell — it’s the whole game!

    • Zacryon@feddit.org
      link
      fedilink
      English
      arrow-up
      7
      ·
      23 days ago

      do to identify AI text — it has lots of tells anyway

      I see what you did there.

  • username_1@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    12
    ·
    24 days ago

    But it doesn’t work. It looks like only the owner of the text generator is able to check if some text is written with this concrete generator (with some probability). I see no use of this technique.

    • fluxx@mander.xyz
      link
      fedilink
      English
      arrow-up
      19
      ·
      24 days ago

      It works for them not scraping their own slop back into training data. I assume that is actually the real purpose of the system. They don’t want to share the key with the public. But they probably will with other llm companies in exchange for theirs.

        • 🌞 Alexander Daychilde 🌞@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          24 days ago

          The article seemed to claim that some models allow outsiders to submit text for detection if I read it right. That seems like a decent way of doing it. If you “open source” the raw data, it means people can do things to try and get around it - same reason most websites don’t reveal their anti-spam techniques.

      • brsrklf@jlai.lu
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        1
        ·
        24 days ago

        That would only work for their own slop though. Anthropic cannot recognize Google’s watermark, only theirs.

        I assumed the goal might be so they could check whether other models have been trained on their output. Like anthropic using that as a “proof” when they start whining again about Chinese “distillation attacks”.

        Of course, since they’re the only ones able to check their watermark, it would be rather shit as evidence anyway. “We’ve run the numbers, and we know you can’t, but trust us, this chatbot is totally copying ours!”

    • brsrklf@jlai.lu
      link
      fedilink
      English
      arrow-up
      6
      ·
      24 days ago

      It’s practically not of any use to end-users. It’s a tool made by the AI company for themselves, to be able to claim a specific text was generated from their model.

      Probability becomes a non-problem the longer the scanned text is.

  • brsrklf@jlai.lu
    link
    fedilink
    English
    arrow-up
    9
    ·
    24 days ago

    I was not sure how any of this worked, and those interactive demos along with the explanation are quite helpful.

    Also a very important point made about this not being a generic AI detector at all, and only being available to the model creator.

  • vane@lemmy.world
    link
    fedilink
    English
    arrow-up
    7
    ·
    23 days ago

    I wonder how many people will be identified as AI because they used AI so much they started constructing sentences like AI.

    • kromem@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      23 days ago

      This particular watermarking would be effectively impossible for a person to end up replicating.

        • kromem@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          21 days ago

          It won’t, unless you normally talk almost exactly like the model in question and then alter your word choice distribution according to the specific secret key entropy.

    • brucethemoose@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      24 days ago

      That’s like 98% of abuse, though. The vast majority of AI abuse is thoughtless, utter laziness.

      Bad actors who actually care about “evading detection” wouldn’t use Claude, anyway.

    • Jestzer@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      24 days ago

      Which, in fairness, there are plenty of. I have seen 2 presentations this week at work where they were 1. Obviously AI-generated and 2. Obviously not proofread or changed at all.

  • Zacryon@feddit.org
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    2
    ·
    23 days ago

    Although this looks like a clever approach, a kind of stochastic key, I do not see how this guarantees to distinguish text written by big babble machines versus humans. Humans also have a certain pattern of writing, a given distribution of how some words are more likely to appear than others. How can one tell them really apart?

    As an indicator, yeah, might be usable. But I wouldn’t read too much into it before seeing results of a study that runs actual tests.

    • kromem@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      23 days ago

      It’s not about the variation of the words, it’s about the variation of the words from the model baseline.

      Like if your word choice was almost the exact same as Claude’s normally, maybe you just talked to them a lot and picked up their phrases like it’s not nothing.

      But if you managed to be almost exactly like Claude and yet varied the possible words exactly according to a hidden entropy key, they’d know it was actually Claude with the SymthID-Text watermarking applied, as no human would end up falling into that statistical bucket.

      • Zacryon@feddit.org
        link
        fedilink
        English
        arrow-up
        1
        ·
        22 days ago

        Yeah, still, I wouldn’t claim “as no human would end up falling into that”, given that it may not be that unlikely to find at least one human who displays similar writing the more humans you involve.

        Until a formal analysis is presented and an experimental study is published, which covers the most important influencing factors, the reliability of this concept is limited.

    • Angry Fuck@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      23 days ago

      I thought the article explained that pretty reasonably on a scale of probability and weight. The longer the text, the more reliable the scoring.

      • Zacryon@feddit.org
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        2
        ·
        22 days ago

        But it does not show a sufficient formal proof and no experimental validation. Many important questions to evaluate the concept are left unanswered, which limits the interpretability and condenses it to “just trust me, bro, it’s a good idea, because I say so”.

        • Angry Fuck@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          22 days ago

          I’m not sure we’ve read the same article. There are literally interactive demonstrations within the page to demonstrate how the concept works.

          • Zacryon@feddit.org
            link
            fedilink
            English
            arrow-up
            1
            ·
            22 days ago

            Interactive demonstrations are not the same as a formal proof or experimental validation. So we shouldn’t attribute more to this technique than the available evidence can really support.

            I found some time to quickly skim through the sources they have listed. And from that it became pretty clear that this is not realiable in detecting LLM generated versus human output in general. Under very tight assumptions specific error rates were reported that appeared rather low. However, these assumptions do not hold in general, even with more text if no relevant signal remains. There is currently no scientifically validated general purpose way of reliably detection.

            More importantly in the context of Claude, the production watermarking scheme is undisclosed. Therefore, the cited experiments on known watermarking schemes can neither establish how reliably text generated by Claude can be detected, nor how reliably the technique described in the article removes the actual watermark.

            It can be treated as an indicator at best, but not as validated proof.

            • Angry Fuck@lemmy.world
              link
              fedilink
              English
              arrow-up
              2
              ·
              22 days ago

              Gish Gallop… if you’re going to start questioning whether the technique clearly demonstrated has validity, then you need to specifically state what your objections are, as opposed to vague statements. For emphasis, the demonstration isn’t on AI detecting, but rather AI watermarking. You wouldn’t use this tool to check if text was written by AI, but rather if the text was written by one singular LLM vs literally everything else.

  • antianarchist@sopuli.xyz
    link
    fedilink
    English
    arrow-up
    6
    arrow-down
    1
    ·
    24 days ago
    • Only the key-holder can check. Your teacher, editor, or favourite “AI detector” website cannot run this test; a genuine check needs the provider’s secret key, or a checking service the provider runs. Google runs an early-access detector portal for SynthID; Anthropic says detection tooling is forthcoming.

    I am not so sure about that. The amounts of words is finite and with enough text, you will see that certain words are used more often, especially in certain combinations. I believe people will brute force this and then create a way to destroy the watermark again.

    • trem@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      2
      ·
      24 days ago

      Yeah, they can’t easily rotate keys, because the text can’t tell you which key was used.

      They could switch to a new key e.g. every month and then just check every previous key during detection. But that would slowly increase the likelihood of false positives, so no idea if that’s a good idea either.

    • Zacryon@feddit.org
      link
      fedilink
      English
      arrow-up
      1
      ·
      23 days ago

      with enough text, you will see that certain words are used more often

      Which is also a thing humans do.

      • antianarchist@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        2
        ·
        23 days ago

        Absolutely! We all basically do fingerprinting. We’re just not really conscious about the key we are using. But with a bit of statistics, you could identify people.

        • Zacryon@feddit.org
          link
          fedilink
          English
          arrow-up
          2
          ·
          22 days ago

          But how reliable? With which guarantees? What are the prerequisites for this to work at all? Telling people from each other apart is one thing, the other is telling them reliably apart from a machine generated text.

  • LedgeDrop@lemmy.zip
    link
    fedilink
    English
    arrow-up
    3
    ·
    24 days ago

    Great article. I wonder if the same markers can be used to detect AI generated code (if you suppressed comments).

    As, code requires a much more rigid syntax, compared to free flowing docs.

    • swicano@programming.dev
      link
      fedilink
      English
      arrow-up
      2
      ·
      24 days ago

      Definitely. It might require significantly more input to gain the same certainty, but all it’s doing is reweighting the possible next tokens before choosing, and code output is still just token output. The rigid syntaxes probably means that the next token probabilities are much more sharply divided (maybe a random sentence the top 1 choice is just 40%, top 3 are 80%, but for a line of code, the top 1 choice might be 90% probability and top 3 hit 99%)

    • jj4211@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      24 days ago

      You still have choices.

      Variable names are pretty free form.

      A switch statement or if/else might be a choice that achieves the same thing. Waffling between them would be highly suspect, since a person isn’t going to be so wishy washy. So you could use structural choices too.

      AI code tends to look more obviously AI than prose anyway.

      • DeadDigger@lemmy.zip
        link
        fedilink
        English
        arrow-up
        1
        ·
        23 days ago

        I mean 50% of the time you just use the name intellij or vs code suggests and the other 48% are slight variances

  • Wispy2891@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    23 days ago

    I wonder if they did this to appease the EU or just to have a way to prove in court that a specific competitor distilled their model using claude

    • Angry Fuck@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      23 days ago

      Sell access to Turnitin and the likes for a small fortune. They are all but required to pay whatever the price is.

  • shoo@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    arrow-down
    1
    ·
    24 days ago

    Unless I missed something, that seems pretty brittle. Wouldn’t any minor editing break it because the watermark is derived from the preceding text? Eg. Find + replace “it is” to “it’s”

    • trem@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      3
      ·
      24 days ago

      The section “4. What editing does to the mark” talks about that. Probably best to look at that illustration again, but basically those edits would interrupt consecutive runs of detectable text, but if a run is long enough, it can still be detected with statistical significance.

      So, it doesn’t have to check the ‘color’ of the words from start to end uninterrupted, but rather can also detect color sequences in the middle of the text.

    • dream_weasel@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      3
      ·
      24 days ago

      Yes you missed at least one whole section including a graphic that shows the breakdown of the watermark with typo fixing, light paraphrasing, moderate and heavy editing.

  • Zarobi@aussie.zone
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    1
    ·
    24 days ago

    What is the statistical likelihood of an individual possessing access to a thesaurus inadvertently precipitating the activation of the artificial-intelligence revelation watermark?

    It’s not like it’s a secret invisible Unicode character flag or something, it’s just a series of word choices. To me, this seems extremely unreliable. It’s only one step removed from those “unreliable A.I. detection tools” that scan for common word choices A.I. uses. You’ve just biased your own A.I. to use specific word choices and then told your own A.I. to check for those words. This doesn’t seem special or interesting to me. There’s still going to be false positives, but now with even more false confidence.

    • Funkt4st1c@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      24 days ago

      Lots of people say my writing looks like AI because i use em dashes (learned about them in 8th grade) and semicolons (6th grade). My only saving grace is my extremely long sentences; AIs tend to have shorter, more poignant sentences with obvious-ish tells once you know what to look for.

      Maybe also helps that i changed keyboards recently so i type weird words like “knkw” instead of “know” and dont always double check my spelling