you are viewing a single comment's thread
view the rest of the comments
[–] 69 points 4 months ago* (12 children)

I kinda want to mirror this to the fediverse with a bot to 1. Make more people see it and 2. Mirror it so when it gets taken down its distributed on here.

Should I do it? Or is that dumb?

  • source
  • hideshow 12 child comments
  • [–] 7 points 4 months ago (1 child)

    Yeah, I'm real torn. On one hand, I immediately want to scrape this site, but I also don't want to beat the site up tying up their bandwidth. There seems to be a parent site db4p.org thats managing mirrors of this site, but I don't see any sort of torrent or archive. If there's something like that, I'd be very inclined to just archive the entire site/database.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 7 points 4 months ago*

    Mmm… such a bot could run once every 24 hours either “visiting the site” and reading the HTML contents. Or using the DB directly if they have an API somewhere.

    Either way it doesn’t cost them much.

  • source
  • parent
  • [–] 2 points 4 months ago*

    wait is that website a fediverse instance actually?

    also if you do mirror it, make sure to do it in an efficient way. for example, some websites offer one large download to archive the whole site, like wikipedia. is less strain on the server than scraping each page individually.

  • source
  • parent
  • [+] 2 points 4 months ago (4 children)
  • [–] 1 point 4 months ago* (3 children)

    I was thinking indeed scraping (or using an API), and when a new entry is made, repost it on here (on a seperate community).

  • source
  • parent
  • hideshow 3 child comments