Hardware I'm running:
- 8/16 AMD CPU
- 32GB system RAM
- 12GB GPU VRAM (AMD)
I'm mainly using these MoE models:
Qwen-3.6-35B-A3B
- Q4_K_M quant quality
- 128k context (conversation length before it compacts)
- Gives me about 260 prefill and 19 token gen speeds
Gemma4-26B-A4B
- Q8 quant quality
- 128k context length
- Gives me about 190 prefill and 14 token gen speeds