Asking ChatGPT to Repeat Words ‘Forever’ Is Now a Terms of Service Violation

misk@sopuli.xyz · 7 months ago

Asking ChatGPT to Repeat Words ‘Forever’ Is Now a Terms of Service Violation

MNByChoice@midwest.social · 7 months ago

Any idea what such things cost the company in terms of computation or electricity?

Daxtron2@startrek.website · 7 months ago

That’s not the reason, it’s because it was seemingly outputting training data (or at least data that looks like it could be training data)

MNByChoice@midwest.social · edit-2 7 months ago

Sure, but this cannot be free.

Edit: oh, are you suggesting it is the normal cost? Nuts, chathpt is not repeating forever.

nickwitha_k (he/him)@lemmy.sdf.org · 7 months ago

I think that they were referring to the exploit that was recently published. Google researchers were able to reliably get the LLM to output training data verbatim, including PII.

To me, this reads as damage control for that. Especially as they are being sued for copyright infringement, which they and their proponents have been claiming is impossible (clearly, they were either wrong or lying).

regbin_@lemmy.world · edit-2 7 months ago

It’s definitely cost. There are other ways to make it generate text that is similar to training data without needing it to endlessly repeat words so I doubt OpenAI cares in that aspect.

Daxtron2@startrek.website · 7 months ago

It doesn’t endlessly repeat, there’s a cap on token generation per request. It absolutely is because of the recent “exploit”

regbin_@lemmy.world · 7 months ago

I don’t think they would care if it didn’t get popular and having thousands of people trying it out, eating up huge amount of compute resources.

It’s a known quirk of LLMs.

kromem@lemmy.world · 7 months ago

You’re correct.

While costs are tracked per token, behind the scenes the longer the response the more it costs to continue generating, so millions of users suddenly thinking they are clever replicating what they read getting it to max output tokens is a substantial increase in underlying costs.

The DeepMind researchers were likely doing that by API call, which they were at least paying for on a per token basis.

And the terms hasn’t been updated to prevent it, they’ve always had this item as prohibited:

Attempt to or assist anyone to reverse engineer, decompile or discover the source code or underlying components of our Services, including our models, algorithms, or systems (except to the extent this restriction is prohibited by applicable law).