view the rest of the comments
Technology
This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.
Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.
Rules:
1: All Lemmy rules apply
2: Do not post low effort posts
3: NEVER post naziped*gore stuff
4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.
5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)
6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist
7: crypto related posts, unless essential, are disallowed
This is for qwen 3.5, not 3.6, but the vram requirements are a little out there for most:
https://willitrunai.com/blog/qwen-3-5-27b-vram-requirements
16 to 55GB...
I bought two RTX 3060 with 12 Gb each for a total of 24Gb VRAM. Sionce the cards are not top shelf anymore for gaming, they can be had used in the 200€ range, so my setup was 400€ over the cost of my existing PC. My PC, while not last gen was already pretty beefy, with a Ryzen 9 and 64 Gb RAM. I think 400€ to have a pretty good local AI setup is not bad at all.
I very much regret getting a 3070, with only 8GB of ram... Only reason I am hesitant about 3060s is that dual cards will use more power, and the economics are already pretty dicey when the hosted models are being sold below cost.
I learnt a few years back that the most important driver in GPU. or any other non upgradable memory device obsolescence, is RAM. I got a banger of a deal a few years on a RX 480, but it was 4Gb, adequate at the time. The card was obsolete way faster than the 8GB. As for the power draw, are you planning to have your AI setup running 24/7? Nvtop (linux GPU monitoring app) reports 50-80% of the cards top power draw. There is also the fact that my data doesn't leave my setup.
At the time, it wasn't clear that VRAM beyond 8GB was beneficial. Most cards at the time had basically capped at 8 for ages, and I wanted performance for ray tracing. In hindsight it was a bad call, but for gaming it made sense.
Not 24/7, I'd just start my PC when I need it, but its still more power than I'd like.
Keeping data local does make sense, which is why I have been toying with it. I can run relatively stepped on gemma4 models in 8GB, so for now, I'm happy, but im definitely looking at a Nvidia P40 as a second card.
I can run both the 27B and 35B version of Qwen3.6 on my 10GB VRAM 3080, and they run ok. You don't need to put the entire model in VRAM, even if it's probably beneficial.
Full quantisation? I've only got a 8GB 3070, but I'll give it a go
Edit: Tried the unsloth/qwen3.6 with llama.CPP, and it failed to allocate a 26GB Vulcan buffer and died. Dunno what magic your using, no luck for me though :(
I run it using LM Studio, which defaults to Q4 quantization, I think. I was able to put about 10 layers on the GPU with 64k token context. That put me at about 9.1 GB VRAM usage, leaving some room for Video playback xD
I'll give LM studio a go, thanks.
I mean that's high but not unreasonable, especially with unified ram machines
and it's even lower with MTP https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF
Little outside my budget personally, but I've definitely been itching to pull the trigger on a nVidia P40