https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
I'm running gwen 3.6 with 131k context window on a 3090, it's fast enough and about as good as pay to play Claude at work.
https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
I'm running gwen 3.6 with 131k context window on a 3090, it's fast enough and about as good as pay to play Claude at work.