i'm running moderately quantized models on 24GB VRAM and getting like 30-40 tokens a second. add a zero to the price and it's still not a lot for a company.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: