▲ 186 ▼ The rest of the internet is bad, like real fucking bad, this place and lemmygrad and trueanon (when they're not saying shitty stuff) are like the only good normal healthy places left online. (hexbear.net) submitted 11 months ago by SorosFootSoldier@hexbear.net to c/chat@hexbear.net 86 comments fedilink hide all child comments Just putting that out there. While we might have struggle sessions over bullshit, the larger internet zeitgeist is putrid and rancid.
[–] m532@lemmygrad.ml 12 points 11 months ago (1 child) In diffusion, this has already been done. Most models that were made after SD1.5 have a "handpicked" input dataset. I guess its because most of SD1.5 input had garbage quality, which transferred over to the output. permalink fedilink source parent hideshow 2 child comments replies: [–] Carl@hexbear.net 8 points 11 months ago* (1 child) I have to check that out at some point, models like Gemini and GPT take up all the space in the room and it's easy to forget that there's others permalink fedilink source parent hideshow 2 child comments replies: [–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent
[–] Carl@hexbear.net 8 points 11 months ago* (1 child) I have to check that out at some point, models like Gemini and GPT take up all the space in the room and it's easy to forget that there's others permalink fedilink source parent hideshow 2 child comments replies: [–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent
[–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent