I'd link to some blog posts about this an example, but the site they're from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?

you are viewing a single comment's thread
view the rest of the comments
[–] 3 points 3 hours ago* (4 children)

Everybody in this thread is talking about what they're doing, but not a single reply has been able to explain why the bad behavior somehow benefits the companies doing it.

I don't believe the only reason is incompetence; there's got to be somehing else to it.

  • source
  • hideshow 4 child comments
  • [–] 3 points 3 hours ago (1 child)

    To lock all knowledge in their models as the onky source, destroy independent websites , destroy physical books after scanning , destroy peoples brains

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 22 minutes ago

    They're not destroying the books "after" scanning 'for the evulz' to deprive the world of them. They're destroying them before scanning because cutting the spines off with a bandsaw lets them drop the stack of pages in a sheet-feed scanner and get the job done a lot faster than using a book scanner.

    Similarly, I'm sure they've got some actual rationale for their crawler behavior that makes sense (at least to them) and isn't a conspiracy theory.

    Nobody is a mustache-twirling villain in their own mind, even if they turn out to be so from everybody else's perspective.

  • source
  • parent