you are viewing a single comment's thread
view the rest of the comments
[–] 7 points 1 month ago (3 children)

https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant

I'm running gwen 3.6 with 131k context window on a 3090, it's fast enough and about as good as pay to play Claude at work.

  • source
  • parent
  • hideshow 3 child comments