I'd link to some blog posts about this an example, but the site they're from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?

you are viewing a single comment's thread
view the rest of the comments
[–] 2 points 1 hour ago* (2 children)

It’s not necessarily the big companies themselves doing it? But, there’s a push right now to scrape as much data as possible to feed the training of the AI models, and if your business model is entirely based on selling that data you don’t really care who you’re going to piss off. So, in turn these new scraper bots basically behave in the same way that attack bots do where they’ll be behind VPNs and switch IPs and stuff if you block them.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 34 minutes ago (1 child)

    But WTF is the point of scraping the same data over and over multiple times a second before it even has a chance to change? That's just a waste of resources even on the scrapers' part, because that bandwidth could be used grabbing some other new page instead!

  • source
  • parent
  • hideshow 1 child comment