you are viewing a single comment's thread
view the rest of the comments
[–] 8 points 1 day ago (3 children)

Think of it like making a xerox copy of a xerox copy. The copy of the copy is always shittier.

Using synthetic data can escalate model collapse, as a model is only as good as its training data, which is partly why these LLM models "hallucinate", having been trained on a wealth of garbage from Reddit.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 5 points 20 hours ago

    I was going to make a snarky comment about

    "How's it thinking then?! I was trained by adults and so forth back generations!"

    Then I looked at the quality of people trained by other humans instead of nature, comparing myself to my ancestors, looking at the society around me.

  • source
  • parent
  • [–] 3 points 1 day ago

    So that means when I publish a nifty FOSS project, I should alongside publish one hundred copies which contain stealthy BS LLM modifications which introduce subtly wrong code (like failing invariants or undefined ehavior in concurrent C++ code)? And all dated back to 2020?

    Got it!

  • source
  • parent