you are viewing a single comment's thread
view the rest of the comments
[–] 2 points 2 weeks ago*

the 6B actives should easily fit into an 8GB card

That’s not how MoE models work. There are many “expert’ models and there is a static router model (dense) which determines which “expert” models to route the tokens through. What you want at a minimum is the dense portion of the model to be on VRAM and all of the weights to be in RAM/VRAM for best performance.

  • source
  • parent