▲ 49 ▼ DeepSeek V4 Pro 0813 Weights Released! (huggingface.co) submitted 1 week ago by SirDimples@programming.dev to c/localllama@sh.itjust.works 8 comments fedilink hide all child comments This is an MoE model with 1.6T-A49B The weights were up briefly then taken down due to some issues in the repo files apparently, now they're back up: ModelScope HuggingFace GGUFs are out as well: Unsloth's DeepSeek V4 Pro 0813 DeepSeek published benchmarks for reference:
[–] JoMiran@lemmy.ml 2 points 1 week ago (2 children) 96gb permalink fedilink source parent hideshow 4 child comments replies: [–] lynx@sh.itjust.works 2 points 5 days ago With Q4 everything below 150B should be fine. You can also run the -Flash variant of this model in Q1, but it is probably not usable. permalink fedilink source parent [–] e0qdk@reddthat.com 2 points 1 week ago If I understand the nature of your hardware correctly, you should be able to run the MoE models like Gemma4 26B-A4B or Qwen3.6 35B-A3B at a high quantization fairly performantly. You could try running some of the dense models (like today's Qwen 3.8 27B) as well, but I expect they'll be pretty slow (judging by my own experience with a unified RAM system that has a Strix Halo APU). Might still be useful for tasks that you can leave running on their own for a long time instead of for interactive chat style interaction though. You've got enough RAM to load larger models, but there hasn't been much released in between the "it fits on a 24GB or 32GB GPU that a gamer might own" and the "oh god you need HOW MUCH RAM!?" scales lately... permalink fedilink source parent
[–] lynx@sh.itjust.works 2 points 5 days ago With Q4 everything below 150B should be fine. You can also run the -Flash variant of this model in Q1, but it is probably not usable. permalink fedilink source parent
[–] e0qdk@reddthat.com 2 points 1 week ago If I understand the nature of your hardware correctly, you should be able to run the MoE models like Gemma4 26B-A4B or Qwen3.6 35B-A3B at a high quantization fairly performantly. You could try running some of the dense models (like today's Qwen 3.8 27B) as well, but I expect they'll be pretty slow (judging by my own experience with a unified RAM system that has a Strix Halo APU). Might still be useful for tasks that you can leave running on their own for a long time instead of for interactive chat style interaction though. You've got enough RAM to load larger models, but there hasn't been much released in between the "it fits on a 24GB or 32GB GPU that a gamer might own" and the "oh god you need HOW MUCH RAM!?" scales lately... permalink fedilink source parent