I'd link to some blog posts about this an example, but the site they're from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?

you are viewing a single comment's thread
view the rest of the comments
[–] 5 points 8 hours ago*

The latest scrape I saw that took down one of my sites just threw shit at the wall until something stuck. It took one second and did 50+ requests mostly for pages that didn’t exist.

  • source