▲ 186 ▼ The rest of the internet is bad, like real fucking bad, this place and lemmygrad and trueanon (when they're not saying shitty stuff) are like the only good normal healthy places left online. (hexbear.net) submitted 11 months ago by SorosFootSoldier@hexbear.net to c/chat@hexbear.net 86 comments fedilink hide all child comments Just putting that out there. While we might have struggle sessions over bullshit, the larger internet zeitgeist is putrid and rancid.
[+] vegeta1@hexbear.net 58 points 11 months ago* (last edited 6 months ago) (1 child) [deleted] permalink fedilink source hideshow 2 child comments replies: [–] Carl@hexbear.net 32 points 11 months ago* (3 children) I feel like eventually we'll see an "evolution" of LLMs where the big innovation will be cutting 90% of the Internet out of the training data without breaking the whole thing. Imagine if LLM output was as dry, neutral, and reliable as the average encyclopedia (yes I know those aren't perfect either but it's an improvement over reddit threads at least). permalink fedilink source parent hideshow 6 child comments replies: [–] EnsignRedshirt@hexbear.net 21 points 11 months ago (1 child) I don’t know if it’ll be framed as an innovation, per se, but that’s going to be the main utility for this technology. Small, focused models that can help you turn a large amount of pre-qualified data into something usable. That would be pretty cool. Wasn’t ever going to be anything more than that, but we’ll have to watch a trillion dollar market bubble pop before people start to narrow their ambitions and actually make something useful out of these things. permalink fedilink source parent hideshow 2 child comments replies: [–] GrouchyGrouse@hexbear.net 14 points 11 months ago It’s such a depressingly stupid time to waste a bunch of information and digital tech so you can have Racist Google instead of what it was 20 years ago. We’re on the cusp of environmental changes brought about by wasting resources. That they built a giant wasteful bubble is the least surprising part of their behavior. permalink fedilink source parent [–] Le_Wokisme@hexbear.net 20 points 11 months ago reddit was actually good for a bunch of how-to kinda shit that would never be in an encyclopedia. the trick is sifting the "hey you might have a carbon monoxide leak" from the "it's cool to throw car batteries into the sea" permalink fedilink source parent [–] m532@lemmygrad.ml 12 points 11 months ago (1 child) In diffusion, this has already been done. Most models that were made after SD1.5 have a "handpicked" input dataset. I guess its because most of SD1.5 input had garbage quality, which transferred over to the output. permalink fedilink source parent hideshow 2 child comments replies: [–] Carl@hexbear.net 8 points 11 months ago* (1 child) I have to check that out at some point, models like Gemini and GPT take up all the space in the room and it's easy to forget that there's others permalink fedilink source parent hideshow 2 child comments replies: [–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent
[–] Carl@hexbear.net 32 points 11 months ago* (3 children) I feel like eventually we'll see an "evolution" of LLMs where the big innovation will be cutting 90% of the Internet out of the training data without breaking the whole thing. Imagine if LLM output was as dry, neutral, and reliable as the average encyclopedia (yes I know those aren't perfect either but it's an improvement over reddit threads at least). permalink fedilink source parent hideshow 6 child comments replies: [–] EnsignRedshirt@hexbear.net 21 points 11 months ago (1 child) I don’t know if it’ll be framed as an innovation, per se, but that’s going to be the main utility for this technology. Small, focused models that can help you turn a large amount of pre-qualified data into something usable. That would be pretty cool. Wasn’t ever going to be anything more than that, but we’ll have to watch a trillion dollar market bubble pop before people start to narrow their ambitions and actually make something useful out of these things. permalink fedilink source parent hideshow 2 child comments replies: [–] GrouchyGrouse@hexbear.net 14 points 11 months ago It’s such a depressingly stupid time to waste a bunch of information and digital tech so you can have Racist Google instead of what it was 20 years ago. We’re on the cusp of environmental changes brought about by wasting resources. That they built a giant wasteful bubble is the least surprising part of their behavior. permalink fedilink source parent [–] Le_Wokisme@hexbear.net 20 points 11 months ago reddit was actually good for a bunch of how-to kinda shit that would never be in an encyclopedia. the trick is sifting the "hey you might have a carbon monoxide leak" from the "it's cool to throw car batteries into the sea" permalink fedilink source parent [–] m532@lemmygrad.ml 12 points 11 months ago (1 child) In diffusion, this has already been done. Most models that were made after SD1.5 have a "handpicked" input dataset. I guess its because most of SD1.5 input had garbage quality, which transferred over to the output. permalink fedilink source parent hideshow 2 child comments replies: [–] Carl@hexbear.net 8 points 11 months ago* (1 child) I have to check that out at some point, models like Gemini and GPT take up all the space in the room and it's easy to forget that there's others permalink fedilink source parent hideshow 2 child comments replies: [–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent
[–] EnsignRedshirt@hexbear.net 21 points 11 months ago (1 child) I don’t know if it’ll be framed as an innovation, per se, but that’s going to be the main utility for this technology. Small, focused models that can help you turn a large amount of pre-qualified data into something usable. That would be pretty cool. Wasn’t ever going to be anything more than that, but we’ll have to watch a trillion dollar market bubble pop before people start to narrow their ambitions and actually make something useful out of these things. permalink fedilink source parent hideshow 2 child comments replies: [–] GrouchyGrouse@hexbear.net 14 points 11 months ago It’s such a depressingly stupid time to waste a bunch of information and digital tech so you can have Racist Google instead of what it was 20 years ago. We’re on the cusp of environmental changes brought about by wasting resources. That they built a giant wasteful bubble is the least surprising part of their behavior. permalink fedilink source parent
[–] GrouchyGrouse@hexbear.net 14 points 11 months ago It’s such a depressingly stupid time to waste a bunch of information and digital tech so you can have Racist Google instead of what it was 20 years ago. We’re on the cusp of environmental changes brought about by wasting resources. That they built a giant wasteful bubble is the least surprising part of their behavior. permalink fedilink source parent
[–] Le_Wokisme@hexbear.net 20 points 11 months ago reddit was actually good for a bunch of how-to kinda shit that would never be in an encyclopedia. the trick is sifting the "hey you might have a carbon monoxide leak" from the "it's cool to throw car batteries into the sea" permalink fedilink source parent
[–] m532@lemmygrad.ml 12 points 11 months ago (1 child) In diffusion, this has already been done. Most models that were made after SD1.5 have a "handpicked" input dataset. I guess its because most of SD1.5 input had garbage quality, which transferred over to the output. permalink fedilink source parent hideshow 2 child comments replies: [–] Carl@hexbear.net 8 points 11 months ago* (1 child) I have to check that out at some point, models like Gemini and GPT take up all the space in the room and it's easy to forget that there's others permalink fedilink source parent hideshow 2 child comments replies: [–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent
[–] Carl@hexbear.net 8 points 11 months ago* (1 child) I have to check that out at some point, models like Gemini and GPT take up all the space in the room and it's easy to forget that there's others permalink fedilink source parent hideshow 2 child comments replies: [–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent
[–] piccolo@hexbear.net 5 points 11 months ago The other person was talking about image generation models, not LLMs. I think that the only LLMs with super curated input sets are tiny and less useful. Unfortunately it takes a lot of data for LLMs to be trained so it's hard to find enough good quality data if you're curating it. permalink fedilink source parent