I'd link to some blog posts about this an example, but the site they're from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host's resources beyond being an ⊛ to webmasters?

you are viewing a single comment's thread
view the rest of the comments
[–] 43 points 6 hours ago*

You are sitting behind a receptionist’s desk with three coworkers. You need to check people into your office for visits.

Because there are four of you, and people are generally good about queuing, your normal operations run smoothly.

Now fifty people walk into the building at once. They don’t really care about checking in, or even visiting the office. They want to know if you have free coffee, a public bathroom, your elevator inspection on file, the date of your last fire inspection, and dozens more inane questions that aren’t what you’re used to handling. They refuse to wait in a line, they’re talking over one another, and if you take more than ten seconds to answer, they walk out of the building but come back to bother you a few minutes later.

Does your office run well, and can you check in a legitimate visitor in a timely manner still?

  • source