▲ 416 ▼ Cara is finally taking action against AI-bro scrapers (thelemmy.club) submitted 3 days ago* (last edited 3 days ago) by destructdisc@lemmy.world to c/fuck_ai@lemmy.world 54 comments fedilink hide all child comments Context: https://blog.cara.app/blog/cara-needs-your-help
[–] AeonFelis@lemmy.world 4 points 2 days ago (1 child) Don't they use Anubis? permalink fedilink source hideshow 2 child comments replies: [–] ruby@lemmy.dbzer0.com 3 points 2 days ago (1 child) i just visited and no, they don't seem to use it, or at least i could access the site without it. but even if they did, anubis is incapable of defending against targeted scraping. permalink fedilink source parent hideshow 2 child comments replies: [–] emmowo@lemmy.world 1 point 2 days ago* (1 child) Anubis can do an okay job... (at least so the whole site isn't being actively DDos'ed) at the cost of compromising user experience a lot by sending harder challenges more often. It kinda sucks that the average user might not have that fast of a CPU, while motivated scrapers might spend tonnes on compute just out of spite. (i am a bit biased on this though) permalink fedilink source parent hideshow 2 child comments replies: [–] ruby@lemmy.dbzer0.com 2 points 1 day ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent
[–] ruby@lemmy.dbzer0.com 3 points 2 days ago (1 child) i just visited and no, they don't seem to use it, or at least i could access the site without it. but even if they did, anubis is incapable of defending against targeted scraping. permalink fedilink source parent hideshow 2 child comments replies: [–] emmowo@lemmy.world 1 point 2 days ago* (1 child) Anubis can do an okay job... (at least so the whole site isn't being actively DDos'ed) at the cost of compromising user experience a lot by sending harder challenges more often. It kinda sucks that the average user might not have that fast of a CPU, while motivated scrapers might spend tonnes on compute just out of spite. (i am a bit biased on this though) permalink fedilink source parent hideshow 2 child comments replies: [–] ruby@lemmy.dbzer0.com 2 points 1 day ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent
[–] emmowo@lemmy.world 1 point 2 days ago* (1 child) Anubis can do an okay job... (at least so the whole site isn't being actively DDos'ed) at the cost of compromising user experience a lot by sending harder challenges more often. It kinda sucks that the average user might not have that fast of a CPU, while motivated scrapers might spend tonnes on compute just out of spite. (i am a bit biased on this though) permalink fedilink source parent hideshow 2 child comments replies: [–] ruby@lemmy.dbzer0.com 2 points 1 day ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent
[–] ruby@lemmy.dbzer0.com 2 points 1 day ago anubis does its thing by making you waste your cpu cycles and then giving you a cookie, if you use that cookie for subsequent requests then you're free to browse the site. it might in theory slow down dumb scrapers that don't save that state and parallelize the mass scraping via multiple ip addresses (since without the cookie they'll have to solve the challenge for each request) but if one's making a targeted attack for one specific site, they don't do any of that. it should be trivial for a bad actor to solve the challenge once and use the cookie, like a regular user would. anubis only slows down scrapers that scrape all sites indiscriminately and don't expect it, once you know it's there then it's easy to bypass (and if you can afford to scrape the entire site you can surely afford to get past anubis once). permalink fedilink source parent