434
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 03 Aug 2026
434 points (91.6% liked)
Technology
86866 readers
4912 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
I'm not sure about the weed argument, can't say anything about it. But first, my personal problem with Ai doesn't matter to this disussion, as I was talking about what people could have an issue with, not about my personal problems. Secondly farming data without respecting the original license, not even linking to the source where it got from is in fact a problem with "its use". That is an unsolved problem lot of people just ignore, which does not make it less of a problem (it makes it a bigger problem). And that is just a few issues I listed, not even talking about the problems of generated content. There are legitimate concerns using Ai, no matter how you frame it, some just choose to ignore it.
Okay. Then the problem that people, not you particularly, have with AI isn’t with the technology itself nor its use, but the methods by which the companies "grow and harvest it", so to speak.
Which I don't even think is true. People have just lost their minds over this. If a company unethically steals data and publishes a book with that data, we don't say that there are legitimate concerns using books.
We would be having a problem with it if every book was regurgitated stolen data. That's kind of the core of how all of these generative AI models work to the point that their creators have argued in court that they wouldn't exist at all if they weren't allowed to steal all the training data.
It can be trained on the exact same free data that we can be trained on. Just like us, it doesn't have to bypass paywalls. I think they're arguing that if I can look at the Wikipedia and tell you that an avalanche in Pakistan killed ten climbers, then AI should be allowed to as well.
I think you need to read up on the difference between "free as in beer" and "free" licenses. Everything on Wikipedia is copyrighted under a GNU license, which requires any use to give credit AND ALSO be licensed under the GPL. Which AI models do not (and cannot) do.
Also, just conceptually... We built a giant, free library (and museum and art gallery and etc.) on the internet, and these parasitic AI companies came in and built a fence around it and started charging admission while massively polluting our cities and towns and gobbling up resources at an incredible rate. Then, even if you pay them their admission fees, the information you get is at best 70% correct. Why would anyone celebrate that?
This is misleading and dishonest. Its not just looking up a single free information. Ai does more than that. Ai will not only scrape and copy the entire article, it will also scrape and copy the entirety of Wikipedia. There are licenses attacked to the Wikipedia article we have to follow to, if we scrape and use it for other purposes. It will also scrape the entire web. That is not the same as just looking up an Avalanche in Pakistan killed ten climbers.
This is weirdly accusatory. I'm not sitting here in a black cape and stovepipe top hat, twirling my mustache.
If this is the way we're going to look at it then I guess we have to throw this technology out the window. It seems like a bad idea to do that though.
BTW so you agree with me then, because the only solution you see here to throw it away. Because what I described is not hallucinated by me, this is how its operate. It was to correct you for your wrong example with looking up a single information from Wikipedia.
Why? Besides that's not possible in the near future, it could be made illegal and then suddenly the biggest companies (the driver behind it) would have to stop. This wouldn't eliminate all Ai tools everywhere, so getting rid of it is probably not possible anymore.
But that wasn't my suggestion either. And we don't have to throw it away, we have to comply. It needs to change.
And this was my point from the very beginning. Using AI isn't bad. The problems that people have with it has nothing to do with the technology itself.
We literally discussed about the problems of Ai the whole time... LLMs have conceptional problems, regardless who uses it. It is the technology that is the problem first, second problem are the companies abusing it even more and then the careless or greedy users too. Its threefold. Saying it has nothing to do with the technology itself is therefore wrong.
I disagree with this.
I agree with this.
The problem with everyone's panic over AI is that they can't think straight. It's like if I say that hammers are fantastic but everyone says that they are horrible because they have conceptual problems, murderous users are hitting people in the head with them, the companies are stealing wood and cutting down the rain forest to make them, etc.; meanwhile a YouTuber is caught using one and he's apologizing and saying he has a problem and his wife and family are chewing him out for using a hammer. It's sheer madness.
There's nothing wrong with using a hammer. If the companies are using enslaved ewoks provided by Lord Vader to make them, that has nothing to do with the technology itself.
Again wrong analogy just to disagree. That is not the conceptional problem with LLMs I talked about. It is a matter of fact that the LLMs are trained on data without respecting their license such as GPL, even worse when generating content. These are factual problems with the concept of LLMs. This is not like a hammer that was designed for a specific purpose, that is then misused as something else. That is not the issue whatsoever here. The design and philosophy of LLMs is deeply flawed in many ways.
The hammer analogy shows either you try to change the topic, or you don't understand the issues I'm talking about. That is completely false analogy to the issues I brought up. Same with a knife or even a game controller. Just because you can misuse it, does not mean its wrongly designed. The user didn't use it properly. That is NOT what I was talking about.
You're going in circles, talking out of both sides of your mouth.
You just come up with analogies that does not address the issues I am talking about. We are talking about the source code the LLMs copy and analyze and then not respect the license. They fail to provide the source code and source material they analyzed, so we could build the Ai models from scratch. And I am not talking about the data after training.
Your hammer analogy (and others) are absolutely not addressing any of these issues. Instead you talk about the end user that can misuse the tool, which is NOT what I am talking about. And now you start attacking personally because you ran out of arguments. Just wasting peoples time.
Off course we say its a legitimate concern if someone unethically steals data and publishes a book with that data... The point of scraping and stealing data without respecting its license, training their software with it and using it commercially (or even non commercially) without giving even credit and not respecting its original license?
No, its not just the methods. I am talking about deep problems with the technology, the methods how companies go with it, and then its usage from company and end user. Ai trained with those data and used makes it impossible to respect the original license, for the user. So in this case, the company using the Ai AND the user of the model violate licenses. Maybe. Maybe not, if the data allows it. But we don't know and CANNOT check and know for sure, that's the problem. Just because you don't understand the issue or don't care does not make it less of a problem.
As an example with software licenses like GPL requires that any derivative works using this source code (such as Linux or whatever my personal program is licensed under GPL) has to be GPL licenses (Open Source) too. That is not debatable. That is how this license work. Now the Ai can be trained on this data, using the source code, and then? See where this goes? That is only one generalized example.
We have a legitimate concern with that particular book, not all books.
The difference is, you know which book you have a concern with. Ai does not operate like that. With using Ai, you don't know which book the data is from and it is not possible for the Ai to tell you that. There are conceptional problems with Ai. Your book example does not work here, because that works differently than Ai.