• FlashMobOfOne@lemmy.world
    link
    fedilink
    arrow-up
    8
    arrow-down
    1
    ·
    18 hours ago

    Think of it like making a xerox copy of a xerox copy. The copy of the copy is always shittier.

    Using synthetic data can escalate model collapse, as a model is only as good as its training data, which is partly why these LLM models “hallucinate”, having been trained on a wealth of garbage from Reddit.

    • Nouvellalia@lemmy.world
      link
      fedilink
      arrow-up
      5
      ·
      14 hours ago

      I was going to make a snarky comment about

      “How’s it thinking then?! I was trained by adults and so forth back generations!”

      Then I looked at the quality of people trained by other humans instead of nature, comparing myself to my ancestors, looking at the society around me.

    • HaraldvonBlauzahn@feddit.org
      link
      fedilink
      arrow-up
      3
      ·
      18 hours ago

      So that means when I publish a nifty FOSS project, I should alongside publish one hundred copies which contain stealthy BS LLM modifications which introduce subtly wrong code (like failing invariants or undefined ehavior in concurrent C++ code)? And all dated back to 2020?

      Got it!