I'd link to some blog posts about this an example, but the site they're from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?

you are viewing a single comment's thread
view the rest of the comments
[–] 6 points 7 hours ago (1 child)

To lock all knowledge in their models as the onky source, destroy independent websites , destroy physical books after scanning , destroy peoples brains

  • source
  • parent
  • hideshow 1 child comment
  • [–] 2 points 5 hours ago

    They're not destroying the books "after" scanning 'for the evulz' to deprive the world of them. They're destroying them before scanning because cutting the spines off with a bandsaw lets them drop the stack of pages in a sheet-feed scanner and get the job done a lot faster than using a book scanner.

    Similarly, I'm sure they've got some actual rationale for their crawler behavior that makes sense (at least to them) and isn't a conspiracy theory.

    Nobody is a mustache-twirling villain in their own mind, even if they turn out to be so from everybody else's perspective.

  • source
  • parent