This news item from a few weeks ago revealed that when users created a "public link" of their Claude chat history, that chat history ended up in Google's search queries.
Aren't these urls supposed to be more or less undiscoverable by virtue of the long string of random characters that they contain?
How did Google discover those urls?
Did these people publish the urls somewhere publically themselves?
Or did Google "fish" them out of their gmail inbox or something like that?
Did these people publish the urls somewhere publically themselves?
They created a public link. Google scraped it because it wasn't told not to.
From the article:
A spokesman for Google made clear to the BBC that the company does not control "what pages are made public on the web," saying instead that action comes from websites.
"We give site owners clear controls to decide whether pages can be crawled or indexed, and we always respect those directives."
As the search indexing of the chat logs is no longer occurring, it is likely Anthropic used available tools to quickly block the chat log links from search results. Google's process for a website owner to block a link is straightforward, but must be initiated by a website owner.
Google was not the only one:
Other search engines like Bing, Brave and Duck Duck Go, through which the Claude chat logs also appeared, were approached for comment.