They destroy the books because it allows them to scan them bette
That's definitely one reason, but the copyright argument is also a part of it. The court explicitly said that their digitization is legal only if does not increase the number of copies of the book.
There's no need to do this legally - they could donate used books. The argument shows that the final product is transformative, so it doesn't matter whether they keep the books or not.
They have to keep the digital copy of the book because they add it to the training corpus for the LLM. Selling or donating the original physical book after doing that would void the fair use argument.
LLMs can reproduce quite long passages of books, though I don't think this changes the argument much, because it's not reliable or useful.
One of the tests that determines if it's fair use or not is whether it can serve as a replacement for the original book. Pirated copies can, which is why they're illegal. A summary like CliffNotes can't. Even if the LLM can reproduce long passages, you can't do that reliably (like you said) and it won't work for all books.
I always thought there was in obvious win where companies doing this could be forced to archive the scan publicly (after some period of time)
I definitely agree with this. I think copyright law needs to be modernized to handle cases like this. I think the AI companies should be allowed to donate the digital copy to a library (like the Internet Archive) while still being allowed to keep their copy in their training corpus.