you are viewing a single comment's thread
view the rest of the comments
[–] 2 points 1 month ago* (1 child)

For your hardware, the VRam is not enough to run 27b but, I’d recommend Qwen 3.5 9b for image / text to text.

And I’m planning to experiment with Qwen 3.8 9b for text to text.

4_k_m quantization is the sweet spot for performance and ram usage.

Also, I find Llama cpp is better than Ollama in terms of performance.

  • source
  • parent
  • hideshow 1 child comment