My question would be if it uses stuff like CommonCrawl as training material. Given their size I assume they used anything incl. unethical training data, but at least admit it?
The only models I ever found that even tried to only resort to ethically obtained data, being FOSS etc. were the tiny ones from PleIAs. And as expected they're completely useless.
So far I concluded that an "ideologically sound" LLM is impossible due to lack of training data. Unless your ideology allows to steal stuff.