With the frustrations from using the usual search engines growing I've been thinking about this a lot lately.

It seems we've poisoned the well by allowing the proliferation of advertising interests to dominate the web.

Like how hard would it be to make your own non-commercial index?

Only human made sites that aren't related to buying, selling, marketing, etc.

Could that be a federated open-source project?

you are viewing a single comment's thread
view the rest of the comments
[–] [S] 2 points 12 hours ago (1 child)

Thanks for providing that!

Not for the faint of heart resource-wise, but doesn't sound impossible for a dedicated group.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 5 points 12 hours ago

    Oops, realized I didn't answer your question about actually crawling, dig into common crawl documentation, they provide a bunch of technical data and stats that show you the scale..2-4billion pages per month

    And note CC just does a sample of the pages it finds. So the more monthly dumps don't contain all of the data afaik

    And the number above are for one of the monthly dumps

    https://commoncrawl.github.io/cc-crawl-statistics/plots/crawlermetrics

  • source
  • parent