My argument is that an LLM here is reading the content for different reasons than a student would. The LLM uses it to generate text and answer user queries, for cash. The student uses it to learn their field of study, and then apply it to make money. The difference is that the student internalizes the concepts, while the LLM internalizes the text. If you used a different book that covered the same content, the LLM would generate different output, but the student would learn the same thing.
I know it's splitting hairs, but I think it's an important point to consider.
My take is that an LLM algorithm can't freely consume any copyrighted work, even if it's been reproduced online with the consent of the author. The company would need the permission of the author for the express purpose of training the AI. If there's a copyright, it should apply.
You have me thinking though about the student comparison. College students pay to attend lectures on material that can be found online or in their textbooks. Wouldn't paying for any copyright material be analogous to this?