Community-driven open-souce LLM

lily33@lemmy.world · 1 year ago

If you give me several paragraphs instead of a single sentence, do you still think it’s impossible to tell?

lily33@lemmy.world · edit-2 1 year ago

I don’t see how that affects my point.

Today’s AI detector can’t tell apart the output of today’s LLM.
Future AI detector WILL be able to tell apart the output of today’s LLM.
Of course, future AI detector won’t be able to tell apart the output of future LLM.

So at any point in time, only recent text could be “contaminated”. The claim that “all text after 2023 is forever contaminated” just isn’t true. Researchers would simply have to be a bit more careful including it.

lily33@lemmy.world · 1 year ago

Not really. If it’s truly impossible to tell the text apart, than it doesn’t really pose a problem for training AI. Otherwise, next-gen AI will be able to tell apart text generated by current gen AI, and it will get filtered out. So only the most recent data will have unfiltered shitty AI-generated stuff, but they don’t train AI on super-recent text anyway.

lily33@lemmy.world · 1 year ago

They don’t redistribute. They learn information about the material they’ve been trained on - not there natural itself*, and can use it to generate material they’ve never seen.

Bigger models seem to memorize some of the material and can infringe, but that’s not really the goal.

lily33@lemmy.world · edit-2 1 year ago

Language models actually do learn things in the sense that: the information encoded in the training model isn’t usually* taken directly from the training data; instead, it’s information that describes the training data, but is new. That’s why it can generate text that’s never appeared in the data.

the bigger models seem to remember some of the data and can reproduce it verbatim; but that’s not really the goal.

lily33@lemmy.world · edit-2 1 year ago

It’s specifically distribution of the work or derivatives that copyright prevents.

So you could make an argument that an LLM that’s memorized the book and can reproduce (parts of) it upon request is infringing. But one that’s merely trained on the book, but hasn’t memorized it, should be fine.

lily33@lemmy.world · edit-2 1 year ago

Why should such a thing be assumed???

lily33@lemmy.world · edit-2 1 year ago

It’s actually a real problem on reddit where people spin up fake users to manipulate votes. Reddit hasn’t published how they detect that exactly, but one way to do that is to look for bad voting patters, like if one account systematically upvotes/downvotes another. But you pretty much can’t without knowing the votes.

lily33@lemmy.world · 1 year ago

True - but it’ll be much easier to detect.

lily33@lemmy.world · edit-2 1 year ago

That last point is completely impossible. Don’t forget that I don’t have to run the official lemmy software on my instance. I can make changes: for example, I can add a feature to my instance like “log every post in a separate, local database before deleting it from lemmy”. Nobody else but me will know this feature exists. Or (to be AGPL compliant) have a separate tool to regularly back up my lemmy database, undoing deletions.

As for the second point: I’d say making local votes private and non-local public will be worse for privacy due to causing confusion.

lily33@lemmy.world · edit-2 1 year ago

I’d go the other way: make these things officially public, so people know they are, and then aren’t taken by surprise.

Private voting can be tricky in a federated setting, because I could have a malicious instance that boosts my posts (I can have it with public votes too, but then it’s easier to detect). Truly private posting history is outright impossible, as you said, due to crawlers.

The way to privacy is to make sure not to dox your account, perhaps alternate 2-3 accounts if it’s really important to you.

lily33@lemmy.world · 1 year ago

Frankly, I think someone should actually do that. Except maybe use open source AI instead of ChatGPT.

The fact is, in a federated setting all this data will be accessible. For example, if lemmy tried to hide who made each vote, and just federate totals, that would allow my malicious instance to report 1M upvotes for my post.

When lemmy tries to hide this data, all this does is instill a false sense of privacy with users. IMHO the best thing is to make all this de facto public data, officially public, so everyone knows and can act accordingly.

As for privacy, I’d say the best thing to do is, keep your account anonymous.

lily33@lemmy.world · 1 year ago

HuggingFace looks to me like it’s a corporation. Like, when I click on “about > join us”, I’m sent to their job offer page.

lily33@lemmy.world · edit-2 1 year ago

Community-driven open-souce LLM

lily33@lemmy.world · 1 year ago

Yea, didn’t watch the video, but had to post exactly this!

lily33@lemmy.world · 1 year ago

But honestly, I think people will do better long term if they have to put in even just a little bit of legwork to find the communities with the right fit, and ignore the rest.

That kinda misses the point, though. For me it’s more about promoting decentralization than it’s about whether people’s reasons to want to join all communities on a topic make sense (they actually can for niche topics). Without a feature like that, I fear people will just all join the largest community on the topic and “centralize” it.

lily33@lemmy.world · 1 year ago

I see it as compensating for disadvantages people have. So, if one student has lower test scores, but achieved them despite going to an underfunded school and having a part-time job, then that student scores are actually more impressive than someone else who scored better, but had private tutors throughout high school. Once you account for people’s disadvantages, you should naturally get more diverse student body.

And of course minority students have disadvantages that should be accounted for. But they don’t affect everyone the same, and racial quotas is a very lazy way to do this. Instead, admissions should look at the individual circumstances of each student.

lily33@lemmy.world · 1 year ago

I, personally, want things to be decentralized. I want to have 100+ technology communities that are all relevant. But for that to be practical, there needs to be a simple mechanism for people to follow the topic “technology”, and get the content of all these 100+ communities merged together (then perhaps manually block some of them that have bad moderation). Unless we have such mechanism, we’ll end up with one main big technology community, and all others will be secondary.

lily33@lemmy.world · edit-2 1 year ago

I’m hoping for two features: Let communities “follow” other communities - so one community’s content also shows up on the other. And let me group communities together on my personal feed, if they don’t want to follow each other for some reason. For now, I stay mostly on the home page, which aggregates everything - but I’d much prefer to be able to browse by topic and still have some aggregation.

lily33@lemmy.world · 1 year ago

Why? Colleges can still give preference to students who live in poor neighborhoods or bad school districts. What’s the problem with that approach?

lily33@lemmy.world · 1 year ago

It’s not really human rights violations that drive US sanctions anyway. There’s plenty of other countries that are rife with human rights violations, but don’t get sanctions. So long as they listen when US interests are concerned, they’re fine.