Check what can you use and at what rate of token per seconds would it be... It has examples of many models and quantization levels. Huge resource!

you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 3 months ago

Interesting, I just have 8GB VRAM unfortunately. So can't run anything particularily useful for mye purpose 😔 The Gemma 4 E4B is quite good, but id like to run the 31B one

  • source