Tried the same thing in Asahi but without macOS’ memory management and access to GPU acceleration, it just wasn’t feasible.
Thank you for sharing this result. I knew Asahi's memory management wasn't as robust (so I got a 24GB RAM M2 unit to overcome this).
For your macOS Ollama implementation are you able to leverage the NPU in the hardware (which I know is also unavailable so far in Asahi)?