LLMs are not book reports. They are not synthesizing information. They're just pulling words based on probability distributions.Those probability distributions are based entirely on what training data has been fed into them.
You can see what this really means in action when you call on them to spit out paragraphs on topics they haven't ingested enough sources for. Their distributions are sparse, and they'll spit out entire chunks of text that are pulled directly from those sources, without citation.
If you write a book report that just reprinted significant swaths of the book, that would be plaigerism, and yes, would 100% be called copyright infringement.
Importantly, though, the copyright infringement for these models does not come at the point where it spits out passages from a copyrighted work. It occurs at the point where the work is copied and used for purposes that fall outside what the work is licensed for. And most people have not licensed their words for billion dollar companies to use them in for-profit products.