OpenAI now tries to hide that ChatGPT was trained on copyrighted books, including J.K. Rowling's Harry Potter series::A new research paper laid out ways in which AI developers should try and avoid showing LLMs have been trained on copyrighted material.

you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 3 years ago (1 child)

Should we distinguish it though? Why shouldn't (and didn't) artists have a say if their art is used to train LLMs? Just like publicly displayed art doesn't provide a permission to copy it and use it in other unspecified purposes, it would be reasonable that the same would apply to AI training.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 3 years ago

    Ah, but that's the thing. Training isn't copying. It's pattern recognition. If you train a model "The dog says woof" and then ask a model "What does the dog say", it's not guaranteed to say "woof".

    Similarly, just because a model was trained on Harry Potter, all that means is it has a good corpus of how the sentences in that book go.

    Thus the distinction. Can I train on a comment section discussing the book?

  • source
  • parent