I'd link to some blog posts about this an example, but the site they're from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?

you are viewing a single comment's thread
view the rest of the comments
[–] 2 points 8 hours ago (1 child)

There's canonical url tag or 3xx redirect so a decent crawler can resolve this and abort duplicate operations.

  • source
  • parent
  • hideshow 1 child comment