I use Oobabooga Textgen WebUI (github) as the model loader, and usually an offline model that runs on my own hardware called Flat Dolphin Maid (because it is a merged model derived from 3 and abbreviated into this name). You can find this model on huggingface.co. I use the GGUF version and 4 bit quantization. It works well with 64GB of system memory, 12th gen i7, and a 16GBV 3080Ti GPU. I'm not sure what the practical limitations really are for minimum hardware. This is a 8×7B mixture of experts model.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments