I'd be very surprised if people weren't already scraping Reddit for this.
post
Why is there nothing on reddit about this lol
Can't wait for chatGPT to call me good sir and tell me I win the internet.
FUCK REDDIT! FUCK U/SPEZ! The Red-exit shall endure, VIVA LA LEMMY!!
Deleted? You mean made unscrapeable. It's exclusive to Reddit licensees.
When spez took away API access, he basically shit on the social contract that offered a fair exchange of free access for the content we fed into reddit. After the API change, there were new terms: there is no contract. There are no terms. If you use reddit now, you are giving away everything you are to be indexed and mangled by statistics. You exist as free labor to statisticians and machines.
You are more than a few cents of bad memes.
I'm going to make the request in the AM that Lemmy should add robots.txt rules to disallow AI crawlers, to at least indicate we're not interested. We need legislation that tells scrapers what they can access.
Can someone more savvy explain why they couldn't also scrape what we all say here?
They can and do, but they want the training models to come from highly moderated sources otherwise every AI chatbot would be spewing the most racist parts of 4chan because people would train it that way as a joke.
If you let AI roam freely across the internet, it would only learn porn, sailor moon, dragon Ball z, and nazi germany.
Anything can, the difference is reddit holds the exclusive rights to user comments on their site, and they've chosen to sell it.
Dick dick pussy cunt cock dick pussy ass shit cunt shit motherfucker shit motherfucker ass tits cunt cock motherfucker shit ass tits motherfucker shit c'mon. Scrape that🔥
Not that I’m against telling Reddit to fuck off in no uncertain terms, but won’t providing this kind of poisoning to AI training just make it more resilient to exactly this kind of thing?
They say it’s $60 million on an annualized basis. I wonder who’d pay that, given that you can probably scrape it for free.
Maybe it’s the AI act in the EU. That might cause trouble in that regard. The US is seeing a lot of rent-seeker PR, too, of course. That might cause some to hedge their bets.
Maybe some people had not realized that yet, but limiting fair use does not just benefit the traditional media corporations but also the likes of Reddit, Facebook, Apple, etc. Making “robots.txt” legally binding would only benefit the tech companies.
This is the best summary I could come up with:
Reddit will let “an unnamed large AI company” have access to its user-generated content platform in a new licensing deal, according to Bloomberg yesterday.
The deal, “worth about $60 million on an annualized basis,” the outlet writes, could still change as the company’s plans to go public are still in the works.
The news also follows an October story that Reddit had threatened to cut off Google and Bing’s search crawlers if it couldn’t make a training data deal with AI companies.
Last year, it successfully stonewalled its way out of the biggest protest in its history after changes to its third-party API access pricing caused developers of the most popular Reddit apps to shut down.
As Bloomberg writes, Reddit’s year-over-year revenue was up by 20 percent by the end of 2023, but it was still $200 million shy of a $1 billion target it had set two years prior.
The company was reportedly advised to seek a $5 billion valuation when it opens up for public investment, which is expected to happen in March.
The original article contains 346 words, the summary contains 175 words. Saved 49%. I'm a bot and I'm open source!
all 31 comments