If you're upset with Sam Altman and ChatGPT, then you can simply say "Sam Altman is a thief, and ChatGPT uses stolen data." Which is infinitely more defendable than "AI is trained on stolen data and is unethical."
The issue is, I'd say the same applies to every model that produces useful outputs. LLaMa, Anthropic, Grok, DeepSeek, whatever else is out there. If you tell somebody that the LLM they're using is unethical, they'll nust go use another convenient corporate model.
And "public data" is not enough for me, because that typically means scraping copyrighted content from public websites, I'm not aware of a model that uses only data with permission (either explicit or granted by the license) that's useful, and the people who need to hear about ethical problems with GenAI especially don't know about that.