This news item from a few weeks ago revealed that when users created a "public link" of their Claude chat history, that chat history ended up in Google's search queries.
Aren't these urls supposed to be more or less undiscoverable by virtue of the long string of random characters that they contain?
How did Google discover those urls?
Did these people publish the urls somewhere publically themselves?
Or did Google "fish" them out of their gmail inbox or something like that?
Ah ok, looks like I didn't read the article carefully enough.
That still leaves two scenarios though:
Claude fucked up, and gave Google a way to index public links of people. After this news broke, they asked Google (and other search engines) to not show those urls anymore.
People fucked up and/or fell victim to Google's the data hunger, which let Google to index their public links, after which Claude asked Google not to index such links.
I guess, in the end, I'm mostly wondering if the long random strings of public links are generally "brute-force-discovered" by crawlers like Google, or whether "discovery" of such links requires some kind of breach.