It may be that no one currently knows exactly what these things are trained on, but it could be determined. If you know the methodology you can figure out what data is being used. The companies involved are going to resist letting anyone find out, but I'm hoping a court case will break that black box open.
One of the many problems with this form of AI is the degree to which we don't know where it's getting its information from. Without that, there is no way to determine the reliability of the results. They can sound perfectly reasonable and be entirely untrue.