while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
So either they don't track what goes in their training data, or they do but their LLM is unable to precisely credit sources. Maybe a bit of both. That's convenient, this way they claim great discoveries without crediting all prior work they're relying on.
LLMs are a great plausible deniability generator. These allow one to plagiarize while saying with a straight face no one knows what source material the result are based on.
