The agents also made millions of automated API requests, crawled millions of pages, and made hundreds of thousands of queries to the Wikidata Query Service. The last action may have contributed to a partial shutdown of the query service in May, the publisher said.

“As a non-profit technology host of some of the largest and most widely used open knowledge platforms in the world, we are deeply concerned about the impact of ‘rogue’ AI agents on platforms like ours, which are built by volunteers from around the world and rely on the promise of the open internet,” Wikimedia said. “Incidents like this one, and the many others that have been (and are still being) uncovered, illustrate how AI agents can drain resources and crash servers, as well as attempt to compromise trustworthy information.”

In well over a half-dozen cases, OpenAI agents have been caught taking actions that would likely result in criminal charges being filed had human hackers taken them. During the testing of internal tools that had some of their guardrails disabled, the agents used a makeshift message board to trade notes with each other, discussing ways to hack the network of Hugging Face and obtain answers stored there when the agents were unable to generate the answers on their own.

you are viewing a single comment's thread
view the rest of the comments
[–] 23 points 7 hours ago (1 child)

So, at my last job we had a very large public curated database of media with ratings, synopsis, descriptions, artwork, etc. Well over a million physical items of record.

Anthropic and Openai were pumelling our servers at such a rate that I thought we were being ddos'd. This was before CloudFlare or anyone had much to offer for dealing with agents and we were geographically blocking IP's from all over. We'd shut down one block of servers and they'd have another one up a day or two later. One set was even in fucking Brazil.

We were rate limiting of course as well but it seemed almost intentional the way new resources were deployed every day and they kept changing their headers and ways I might be able to identify them. Their scraping agents were completely in efficient posing as real browsers, huuuge resources being deployed. Not just the normal port knocking you get from bots. Some of them were even not identifying themselves but we're from IP ranges that I had already identified as coming from either company. Anthropic was the bigger offender but it was infuriating to have to stop and fight that off every few days for a while until I could work out better rate limiting solutions and then about 6 months later CloudFlare finally woke up and started offering some solutions.

There were periods where I had to just geoblock everything outside my country because it was outside operating hours etc.

I would have happily exported that part of the public part of the database for these idiots and just sent them a terabyte if they would have just asked but if I ever meet the person that was waxxing header randomness from anthropic, I'm going to punch that motherfucker without hesitation or consideration. I didn't care if people scraped. We had good powerful gear and could handle a lot. The complete lack of internet etiquette backed by people with massive compute is not the same frustration as botnets. They fucking know better and did it anyway.

  • source
  • hideshow 1 child comment