this post was submitted on 29 Sep 2023
439 points (93.5% liked)

Technology

59232 readers
3111 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
 

Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

you are viewing a single comment's thread
view the rest of the comments
[–] [email protected] 2 points 1 year ago

What about my Reddit history?

Arguably there's more of my text there that was used to train these LLMs than most authors in that list.

The comment elsewhere in this thread about models built on broad public data needing to be public in turn is a salient one.

IP laws were designed to foster innovation, not hold it back.

I'd much rather see a world where we have open access models trained broadly and accelerating us towards greener pastures than one where book publishers get a few extra cents from less capable closed models that take longer for us to reach the heyday where LLMs can do things like review the past 20 years of cancer research in order to identify promising trends in allocation of future resources.

OpenAI should probably rightfully be dinged for downloading copyrighted media the same way any average user would be sued when caught doing the same.

But the popular arguments these days for making training infringement are ass backwards and a slippery slope to a far more dystopian future than the alternative.