▲ 973 ▼ Managers (thelemmy.club) submitted 4 months ago* by inari@piefed.zip to c/whitepeopletwitter@sh.itjust.works 180 comments fedilink hide all child comments
[–] Valmond@lemmy.dbzer0.com 3 points 4 months ago (36 children) What kind of hardware would be needed to run such a beast? permalink fedilink source parent hideshow 36 child comments replies: [–] lime@feddit.nu 5 points 4 months ago (35 children) a 128GB framework desktop could do that job. it's increased a bit in price since i last looked at it but €4500 isn't that much for a company. permalink fedilink source parent hideshow 35 child comments replies: [–] theunknownmuncher@lemmy.world 3 points 4 months ago (34 children) Maybe to serve an aggressively quantized model to one very patient user. permalink fedilink source parent hideshow 34 child comments replies: [–] lime@feddit.nu 3 points 4 months ago (33 children) i'm running moderately quantized models on 24GB VRAM and getting like 30-40 tokens a second. add a zero to the price and it's still not a lot for a company. permalink fedilink source parent hideshow 33 child comments replies: [–] theunknownmuncher@lemmy.world 3 points 4 months ago* (last edited 4 months ago) (31 children) Sure, but you're running a very small model compared to what we are talking about. GLM-5.1 is over 200GB even when quantizied to 1-bit. Kimi K2.6 is even bigger. A framework desktop cannot run either of these. Qwen3.6 is significantly smaller and the model weights could fit, but consider the KV-cache you'd need for all of the company's users, and the throughput required to serve them all. You're right that it is within reach for a company but framework desktop makes zero sense for this permalink fedilink source parent hideshow 31 child comments replies: [–] lime@feddit.nu 2 points 4 months ago* (30 children) isn't qwen like 40-50GB? that could work i think. performance is okay even quantised down to 10. permalink fedilink source parent hideshow 30 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent [–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent [–] Evotech@lemmy.world 1 point 4 months ago Yes. It will probably work for 1-2 users at peak. permalink fedilink source parent
[–] lime@feddit.nu 5 points 4 months ago (35 children) a 128GB framework desktop could do that job. it's increased a bit in price since i last looked at it but €4500 isn't that much for a company. permalink fedilink source parent hideshow 35 child comments replies: [–] theunknownmuncher@lemmy.world 3 points 4 months ago (34 children) Maybe to serve an aggressively quantized model to one very patient user. permalink fedilink source parent hideshow 34 child comments replies: [–] lime@feddit.nu 3 points 4 months ago (33 children) i'm running moderately quantized models on 24GB VRAM and getting like 30-40 tokens a second. add a zero to the price and it's still not a lot for a company. permalink fedilink source parent hideshow 33 child comments replies: [–] theunknownmuncher@lemmy.world 3 points 4 months ago* (last edited 4 months ago) (31 children) Sure, but you're running a very small model compared to what we are talking about. GLM-5.1 is over 200GB even when quantizied to 1-bit. Kimi K2.6 is even bigger. A framework desktop cannot run either of these. Qwen3.6 is significantly smaller and the model weights could fit, but consider the KV-cache you'd need for all of the company's users, and the throughput required to serve them all. You're right that it is within reach for a company but framework desktop makes zero sense for this permalink fedilink source parent hideshow 31 child comments replies: [–] lime@feddit.nu 2 points 4 months ago* (30 children) isn't qwen like 40-50GB? that could work i think. performance is okay even quantised down to 10. permalink fedilink source parent hideshow 30 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent [–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent [–] Evotech@lemmy.world 1 point 4 months ago Yes. It will probably work for 1-2 users at peak. permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 3 points 4 months ago (34 children) Maybe to serve an aggressively quantized model to one very patient user. permalink fedilink source parent hideshow 34 child comments replies: [–] lime@feddit.nu 3 points 4 months ago (33 children) i'm running moderately quantized models on 24GB VRAM and getting like 30-40 tokens a second. add a zero to the price and it's still not a lot for a company. permalink fedilink source parent hideshow 33 child comments replies: [–] theunknownmuncher@lemmy.world 3 points 4 months ago* (last edited 4 months ago) (31 children) Sure, but you're running a very small model compared to what we are talking about. GLM-5.1 is over 200GB even when quantizied to 1-bit. Kimi K2.6 is even bigger. A framework desktop cannot run either of these. Qwen3.6 is significantly smaller and the model weights could fit, but consider the KV-cache you'd need for all of the company's users, and the throughput required to serve them all. You're right that it is within reach for a company but framework desktop makes zero sense for this permalink fedilink source parent hideshow 31 child comments replies: [–] lime@feddit.nu 2 points 4 months ago* (30 children) isn't qwen like 40-50GB? that could work i think. performance is okay even quantised down to 10. permalink fedilink source parent hideshow 30 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent [–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent [–] Evotech@lemmy.world 1 point 4 months ago Yes. It will probably work for 1-2 users at peak. permalink fedilink source parent
[–] lime@feddit.nu 3 points 4 months ago (33 children) i'm running moderately quantized models on 24GB VRAM and getting like 30-40 tokens a second. add a zero to the price and it's still not a lot for a company. permalink fedilink source parent hideshow 33 child comments replies: [–] theunknownmuncher@lemmy.world 3 points 4 months ago* (last edited 4 months ago) (31 children) Sure, but you're running a very small model compared to what we are talking about. GLM-5.1 is over 200GB even when quantizied to 1-bit. Kimi K2.6 is even bigger. A framework desktop cannot run either of these. Qwen3.6 is significantly smaller and the model weights could fit, but consider the KV-cache you'd need for all of the company's users, and the throughput required to serve them all. You're right that it is within reach for a company but framework desktop makes zero sense for this permalink fedilink source parent hideshow 31 child comments replies: [–] lime@feddit.nu 2 points 4 months ago* (30 children) isn't qwen like 40-50GB? that could work i think. performance is okay even quantised down to 10. permalink fedilink source parent hideshow 30 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent [–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent [–] Evotech@lemmy.world 1 point 4 months ago Yes. It will probably work for 1-2 users at peak. permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 3 points 4 months ago* (last edited 4 months ago) (31 children) Sure, but you're running a very small model compared to what we are talking about. GLM-5.1 is over 200GB even when quantizied to 1-bit. Kimi K2.6 is even bigger. A framework desktop cannot run either of these. Qwen3.6 is significantly smaller and the model weights could fit, but consider the KV-cache you'd need for all of the company's users, and the throughput required to serve them all. You're right that it is within reach for a company but framework desktop makes zero sense for this permalink fedilink source parent hideshow 31 child comments replies: [–] lime@feddit.nu 2 points 4 months ago* (30 children) isn't qwen like 40-50GB? that could work i think. performance is okay even quantised down to 10. permalink fedilink source parent hideshow 30 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent [–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] lime@feddit.nu 2 points 4 months ago* (30 children) isn't qwen like 40-50GB? that could work i think. performance is okay even quantised down to 10. permalink fedilink source parent hideshow 30 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent [–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] Evotech@lemmy.world 1 point 4 months ago (9 children) And then add 200k context on top And then add hundred of users needing to do things in paralell permalink fedilink source parent hideshow 9 child comments replies: [–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent [–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] lime@feddit.nu 1 point 4 months ago nobody said anything about it being a large company :P anyway, seems the framework is hampered by a slow gpu so the memory issues are apparently moot. permalink fedilink source parent
[–] boonhet@sopuli.xyz 1 point 4 months ago (7 children) If it's a large enough company to have hundreds of users, it can afford several beefy machines tbh permalink fedilink source parent hideshow 7 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] Evotech@lemmy.world 1 point 4 months ago (6 children) It's a capex and that type of hardware needs to be replaced every 3 years minimum and you need people to set it up and maintain a cluster. And it's not straight forward. You are never going to get that approved without a serious business case. Claude on the other end is a opex and much easier to just try out and then build a solution on it Not saying it doesn't happen but it's not as easy as people make it sound like permalink fedilink source parent hideshow 6 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] boonhet@sopuli.xyz 1 point 4 months ago (5 children) It's 3 years if you're trying to be competitive on frontier models and generally capex is preferred to opex because opex never ends I don't think anyone's building a cluster for their business right now, but one single rack after Claude gets rid of their subscription options? Might be a good deal. permalink fedilink source parent hideshow 5 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] Evotech@lemmy.world 1 point 4 months ago (4 children) Capex never ends either if it's hardware. Also you need opex to run it permalink fedilink source parent hideshow 4 child comments replies: [–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] boonhet@sopuli.xyz 1 point 4 months ago (3 children) 400k on a DGX node starts seeming like a great deal when your employees each start using a few hundred dollars worth of Claude tokens every month. That one node can handle a lot of users depending on the model used. It's an expense once every maybe 5 or 6 years in reality and you don't need to hire new people, you just give your existing sysadmins some extra work. They'll complain, but they'll still do it. Of course the sensible alternative is to use a decent model off openrouter for peanuts but then you're sending all your sensitive business secrets to China which is even worse than sharing them with a US AI company. And people WILL be sharing secrets lol permalink fedilink source parent hideshow 3 child comments replies: [–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] Evotech@lemmy.world 1 point 4 months ago (2 children) If only it was that simple permalink fedilink source parent hideshow 2 child comments replies: [–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] boonhet@sopuli.xyz 1 point 3 months ago (1 child) You don't have to run Claude Opus for it to be useful lol permalink fedilink source parent hideshow 1 child comment replies: [–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[–] Evotech@lemmy.world 1 point 3 months ago It's always going to be second rate. And you'll have to defend that permalink fedilink source parent
[+] AtHeartEngineer@lemmy.world 0 points 4 months ago (12 children) [deleted] permalink fedilink source parent hideshow 12 child comments replies: [–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent
[–] lime@feddit.nu 1 point 4 months ago* we were talking about 3.6. deepseek distilled is an alternative that works on more modest hardware. and i'm not really interested in what claude and chatgpt, mistral and the others are doing, i would never tuch those models with a ten foot pole. if i can't run it it does not get run. permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago* (10 children) Qwen3.6 27b beats Claude Opus 4.5 in most benchmarks. Qwen3.6 35b beats Opus 4.5 in a few specific benchmarks, but most benchmarks have Opus 4.5 beating Qwen3.6 35b, although there is not a big gap between Opus 4.5 and Qwen3.6 27b or 35b either way. permalink fedilink source parent hideshow 10 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent
[+] AtHeartEngineer@lemmy.world 0 points 4 months ago (9 children) [deleted] permalink fedilink source parent hideshow 9 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago (8 children) https://github.com/QwenLM/Qwen3.6#benchmarks permalink fedilink source parent hideshow 8 child comments replies: [+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent
[+] AtHeartEngineer@lemmy.world 0 points 4 months ago* (last edited 4 months ago) (7 children) [deleted] permalink fedilink source parent hideshow 7 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent [–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago* (5 children) "I don’t think any of that is true. show me data" is shown data "I won't accept that data!" Lol. Lmao even. Yeah, I'm not going to play this game of trying to anticipate which numbers you're willing to accept and which you aren't. You have just as equal access to a search engine as I have. All of the results I have seen align with the numbers that Qwen released and are well within margins of error. This model's release caused such a stir and was a big deal due to the fact that it reproducibly meets or beats Claude Opus 4.5 while being locally runnable. If you won't believe it, okay, I don't care. 🤷 permalink fedilink source parent hideshow 5 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent [+] AtHeartEngineer@lemmy.world 1 point 4 months ago [deleted] permalink fedilink source parent
[+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 0 points 4 months ago* It's not like the Qwen team hasn't already built a lot of trust with the community. They've never been misleading with previous releases, the "marketing material" (🙄) is for a free product, so they have no incentive to lie, and it would be extra stupid because anyone can run the benchmarks and verify their numbers independently anyway. What would be the point? permalink fedilink source parent
[–] Jiral@lemmy.org 0 points 4 months ago* (6 children) At Q8 it is around 35-40GB I think + memory for required context. I have a Framework desktop. It gets you you around 6t/s. Not suitable for professional use but for personal use I think it is fine. I do prefer Gemma 4 though, but that comes with similar reqirements. permalink fedilink source parent hideshow 6 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] lime@feddit.nu 1 point 4 months ago (5 children) huh, i thought that ryzen ai thing would perform better than that. my 7900xtx regularly gets 30+tps with qwen, up to hundreds with more compressed models. permalink fedilink source parent hideshow 5 child comments replies: [–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] Jiral@lemmy.org 2 points 4 months ago* (4 children) My system runs at 100W TDP though. That is maybe 140W at the power outlet, incl. monitor and everything. This is also the dense 27B model at Q8. But yeah, it is not terribly fast. I think the best use case is on MoE models. GPT-OSS-120B runs on it for example and at 50T/s speed is not a n issue anymore either. (I could get it to run even on just 64GB but the new llama.cpp might need a tiny bit more memory which pushed it just across the limit. yeah I know, for seriously using it you'd need the 128GB version) permalink fedilink source parent hideshow 4 child comments replies: [–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] lime@feddit.nu 2 points 4 months ago (3 children) that's fair, i'm at like 7x the power. the gpu alone easily pulls 350-400W and the rest of the system isn't exactly running lean either. ...man now i really want more vram. permalink fedilink source parent hideshow 3 child comments replies: [–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] Jiral@lemmy.org 1 point 4 months ago* (last edited 4 months ago) (2 children) Yes I think Strix Halo makes sense when low power use is a requirement. I built a custom fanless Strix Halo system for the fun of it and I guess there aren't too many out there running Gemma 4 31B Q8 without a single fan, anywhere. And for MoE models that need 60-80GB + context it is perfect. Those are decently fast then as well. PS: If VRAM is all you care about the maxed out Mac Studio is fascinating. 512GB unified memory for around 10K EUR (pre crazy bubble prices) That should be able to run pretty large MoE models but dense models of that size would probably run glacially. permalink fedilink source parent hideshow 2 child comments replies: [–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] lime@feddit.nu 1 point 4 months ago (1 child) i'm not buying any hardware for the forseeable future :P it's all just wishful thinking at the moment. but unified memory architectureis probably going to become more common so maybe in five years when some new motherboard standard becomes the norm... permalink fedilink source parent hideshow 1 child comment replies: [–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] Jiral@lemmy.org 1 point 4 months ago* I fully understand. ;) Buying hardware now means you'd be either crazy or desperate. permalink fedilink source parent
[–] Evotech@lemmy.world 1 point 4 months ago Yes. It will probably work for 1-2 users at peak. permalink fedilink source parent