LLMs use a new way of data processing (focused on language rules) which is what brought the boost. Not more training data. More training data is needed for an LLM to reach "maturity" but it also increases cost due to resource usage. And current LLMs have already absorbed basically all of the internet and most books, and since the internet is now full of slop, more unpoisoned training data does not exist. That's why more does not help anymore.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: