The issue is, I'd say the same applies to every model that produces useful outputs.
That you know of. Your lack of awareness is not an indication on the stance of the technology itself. If you have a problem with Grok, as I do, then condemn Grok, Twitter, and that pathetic man child that owns them.
If you tell somebody that the LLM they're using is unethical, they'll nust go use another convenient corporate model.
Yeah, so actively fight against these toxic corporate activities. Fuck OpenAI. Fuck Google. Fuck Twitter. But the instant negative reaction to their shared tool is getting quite a few additional people caught in the blast.
And "public data" is not enough for me, because that typically means scraping copyrighted content from public websites, I'm not aware of a model that uses only data with permission
Again, this is an expression of your ignorance, not reality.
OLMo 2 was trained on Wikipedia and other fully public forums. Its training sources and data is fully accessible and open.
GPT-NeoX is an untrained model that you can train yourself. It's literally just the foundations anybody could use.
Pythia is fully open source, and one of its training sources was GitHub. Which does genuinely bring in to question whether or not you can use GitHub to train AI. Personally, I don't see how branching a repository is any different than using the code to train a model. But the current anti-AI trend has people EXTREMELY sensitive to this concept.