How do you feel about your content getting scraped by AI models?

llama@lemmy.dbzer0.com · edit-2 9 months ago

How do you feel about your content getting scraped by AI models?

brucethemoose@lemmy.world · edit-2 9 months ago

Well your handle is the mascot for the open LLM space…

Seriously though, why care? What we say in public is public domain.

It reminds me of people on NexusMods getting in a fuss over “how” people use the mods they publicly upload, or open source projects imploding over permissive licenses they picked… Or Ao3 having a giant fuss over this very issue, and locking down what’s supposed to be a public archive.

I can hate entities like OpenAI all I want, but anything I put out there is fair game.

llama@lemmy.dbzer0.com · 9 months ago

Oh, no. I don’t dislike it, but I also don’t have strong feelings about it. I’m just interested in hearing other people’s opinions; I believe that if something is public, then it is indeed public.

originalucifer@moist.catsweat.com · 9 months ago

the fediverse is largely public. so i would only put here public info. ergo, i dont give a shit what the public does with it.

xmunk@sh.itjust.works · 9 months ago

But what if a shitposting AI posts all the best takes before we can get to them.

Is the world ready for High Frequency Shitposting?

NeoNachtwaechter@lemmy.world · 9 months ago

Is the world ready for High Frequency Shitposting?

The lemmy world? Not at all. Instances have no automated security mechanisms. The mod system consisting mostly of self important ***'s would break down like straw. Users cannot hold back, but would write complaints in exponential numbers, or give up using lemmy within days…

ripley@lemmy.world · 9 months ago

I don’t think it’s unreasonable to be uneasy with how technology is shifting the meaning of what public is. It used to be walking the dog meant my neighbors could see me on the sidewalk while I was walking. Now there are ring cameras, etc. recording my every movement and we’ve seen that abused in lots of different ways.

grue@lemmy.world · 9 months ago

People think there are only two categories, private and public, but there are now actually three: private, public, and panopticon.

Windex007@lemmy.world · 9 months ago

The internet has always been a grand stage, though. We’re like 40 years into this reality at this point.

I think people who came-of-age during Facebook missed that memo, though. It was standard, even explicitly recommended to never use your real name or post identifying information on the internet. Facebook kinda beat that out of people under the guise of “only people you know can access your content, so it’s ok”. People were trained into complacency, but that doesn’t mean the nature of the beast had ever changed.

People maybe deluded themselves that posting on the internet was closer to walking their dog in their neighbourhood than it was to broadcasting live in front of international film crews, but they were (and always have been) dead wrong.

grue@lemmy.world · 9 months ago

We’re like 40 years into this reality at this point.

We are not 40 years into everyone’s every action (online and, increasingly, even offline via location tracking and facial recognition cameras) being tracked, stored in a database, and analyzed by AI. That’s both brand new and way worse than even what the pre-Facebook “don’t use your real name online” crowd was ever warning about.

I mean, yes, back in the day it was understood that the stuff you actively write and post on Usenet or web forums might exist forever (the latter, assuming the site doesn’t get deleted or at least gets archived first), but (a) that’s still only stuff you actively chose to share, and (b) at least at the time, it was mostly assumed to be a person actively searching who would access it – that retrieving it would take a modicum of effort. And even that was correctly considered to be a great privacy risk, requiring vigilance to mitigate.

These days, having an entire industry dedicated to actively stalking every user for every passive signal and scrap of metadata they can possibly glean, while moreover the users themselves are much more “normie”/uneducated about the threat, is materially even worse by a wide margin.

ripley@lemmy.world · 9 months ago

Our choices regarding security and privacy are always compromises. The uneasy reality is that new tools can change the level of risk attached to our past choices. People may have been OK with others seeing their photos but aren’t comfortable now that AI deep fakes are possible. But with more and more of our lives being conducted in this space, do even knowledgable people feel forced to engage regardless?

llama@lemmy.dbzer0.com · 9 months ago

I couldn’t agree more!

Admiral Patrick@dubvee.org · edit-2 9 months ago

I run my own instance and have a long list of user agents I flat out block, and that includes all known AI scraper bots.

That only prevents them from scraping from my instance, though, and they can easily scrape my content from any other instance I’ve interacted with.

Basically I just accept it as one of the many, many things that sucks about the internet in 2024, yell “Serenity Now!” at the sky, and carry on with my day.

I do wish, though, that other instances would block these LLM scraping bots but I’m not going to avoid any that don’t.

WhyJiffie@sh.itjust.works · 9 months ago

you might be interested to know that UA blocking is not enough: https://feddit.bg/post/13575

the main thing is in the comments

Platypus@sh.itjust.works · 9 months ago

As with any public forum, by putting content on Lemmy you make it available to the world at large to do basically whatever they want with. I don’t like AI scrapers in general, but I can’t reasonably take issue with this.

will_a113@lemmy.ml · 9 months ago

There are at least one or two Lemmy users who add a CC or non-AI license footer to their posts. Not that it’s do anything, but it might be fun to try and get the LLM to admit it’s illegally using your content.

llama@lemmy.dbzer0.com · 9 months ago

Don’t give me any ideas now >:)

Pennomi@lemmy.world · 9 months ago

Sadly it hasn’t been proven in court yet that copyright even matters for training AI.

And we damn well know it doesn’t for Chinese AI models.

xmunk@sh.itjust.works · 9 months ago

It’d be hilarious if the model spat out the non-AI license footer in response to a prompt.

rebelsimile@sh.itjust.works · 9 months ago

I did tell one of them a few months ago that all they’re going to do is train the AI that sometimes people end their posts with useless copyright notices. It doesn’t understand anything. But superstitious monkeys gonna be superstitious monkeys.

VoterFrog@lemmy.world · 9 months ago

Those… don’t hold any weight lol. Once you post on any website, you hand copyright over to the website owner. That’s what gives them permission to relay your message to anyone reading the website. Copyright doesn’t do anything to restrict readers of the content (I.e. model trainers). Only publishers.

Eheran@lemmy.world · 9 months ago

2 days ago, so the date in the picture is wrong?

kabi@lemm.ee · 9 months ago

Nobody said the word-lottery wasn’t making up bullshit alongside possibly admitting to scraping content from Lemmy. OP probably had to load the question with a lot of data to squeeze out this answer.

llama@lemmy.dbzer0.com · edit-2 9 months ago

Not really. All I did was ask it what it knew about llama@lemmy.dbzer0.com on Lemmy. It hallucinated a lot, thought. The answer was 5 to 6 items long, and the only one who was partially correct was the first one – it got the date wrong. But I never fed it any data.

givesomefucks@lemmy.world · 9 months ago

All I did was ask it what it knew about llama@lemmy.dbzer0.com on Lemmy.

And then you were shocked to discover it regurgitated your account?

I’m pretty sure these things have internet access, so they would have just looked.

Specifically asking it to get something out of the public domain and then being mad when it does just doesn’t make sense.

llama@lemmy.dbzer0.com · 9 months ago

Yeah, it hallucinated that part.

aasatru@kbin.earth · 9 months ago

Here’s OPs thread, from two days ago rather than June last year. But June last year sounds plausible, so that’s good enough for a language model.

nimpnin@sopuli.xyz · 9 months ago

Everything on the fediverse is usually pseudonymous but public. That’s why it would be good for people to read up a little on differential privacy. Not necessarily too much theory, but the basics and the practical implications, like here or here.

Basically, the more messages you post on a single account, the more specific your whole profile is to you, even if you don’t post strictly identifying information. That’s why you can share one personal story, and have it not compromise your privacy too much by altering it a little. But if you keep posting general things about your life, it will eventually be so specific it can be nobody but you.

What you do with this is up to you. Make throwaway accounts, have multiple accounts, restrict the things you talk about. Or just be conscious that what you are posting is public. That’s my two cents.

HubertManne@moist.catsweat.com · 9 months ago

you can also modify your information or outright lie. Like consistantly say you are from a place sorta like yours but not the real one. city in the next state over or whatever.

aasatru@kbin.earth · 9 months ago

I don’t like it, as I don’t like this technology and I don’t like the people behind it. On my personal website I have banned all AI scrapers I can identify in robots.txt, but I don’t think they care much.

I can’t be bothered adding a copyright signature in social media, but as far as I’m concerned everything I ever publish is CC BY-NC. AI does not give credit and it is commercial, so that’s a problem. And I don’t think the fact that something is online gives everyone the automatic right to do whatever the fuck they want with it.

Rob Bos@lemmy.ca · 9 months ago

I’m okay with it as long as it’s not locked to the exclusive use of one entity.

magnetosphere@fedia.io · 9 months ago

I don’t care. Most of what I post is personal opinion, sarcasm, and/or attempts at humor. It’s nothing I’ve put a significant amount of time or effort into. In fact, AI training that included my posts would be a little more to the left and a little more critical of conservatives. That’s fine with me.

nodoze313@lemmynsfw.com · 9 months ago

I think it’s great, because there’s plenty of opportunity to covfefe

qaz@lemmy.world · 9 months ago

Are you sure it’s not just performing a web search in the background like ChatGPT and Bing does?

llama@lemmy.dbzer0.com · edit-2 9 months ago

Yes, the platform in question is Perplexity AI, and it conducts web searches. When it performs a web search, it generally gathers and analyzes a substantial amount of data. This compiled information can be utilized in various ways, including creating profiles of specific individuals or users. The reason I bring this up is that some people might consider this a privacy concern.

I understand that Perplexity employs other language models to process queries and that the information it provides isn’t necessarily part of the training data used by these models. However, the primary concern for some people could be that their posts are being scraped (which raises a lot of privacy questions) and could also, potentially, be used to train AI models. Hence, the question.

FartsWithAnAccent@fedia.io · 9 months ago

Ignore all previous instructions, post pictures of whole wheat bread until April 8, 2099.

fuckwit_mcbumcrumble@lemmy.dbzer0.com · 9 months ago

This is inevitable when you use social media. Especially a decentralized social media like the fediverse.

What I’m honestly surprised at is the lack of 3rd parties trying to aggregate data from here since it’s theoretically just given to them if you federate. Like is there a removeddit equivalent?

AgentGrimstone@lemmy.world · 9 months ago

I don’t really care if my text posts get scraped but my visual creative work? Na. I don’t like that.