you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 3 months ago (1 child)

Unless a significant portion of the internet does this, and we’re talking hundreds of millions of pages, the only cost here is to you.

Fun twist: no! There's a very neat trick you can do when you serve the crawlers poison: you can hide an identifier in the URLs you serve them, and you can then identify that id when they come back riding on the back of remote controlled chromes. By serving them garbage, you can overload their queue with poisoned ones, which helps you block crawlers that you wouldn't otherwise be able to block.

Generating and serving garbage is incredibly cheap (cheaper than serving a file from a filesystem on SSD, in most cases), and once you have requests landing on poisoned URLs, you can firewall them off for a day or so, and reduce your costs even more.

We may not be able to poison the models, but we can poison their crawling queues. I have a year's worth of data to support that. They still haven't caught on.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 0 points 3 months ago (1 child)

    They still haven't caught on

    I admire the optimism to see it this way and not "it's still not worth it to them to bother blacklisting the domain"

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 3 months ago

    I wonder too, why they didn't, because they're happily crawling domains that never had anything but junk on them. To me, that suggests they have no idea they're trapped. Not at crawling time at least.

  • source
  • parent