you are viewing a single comment's thread
view the rest of the comments
[–] 35 points 1 month ago* (19 children)

Local 27b models are good enough for most tasks.

Can’t wait to buy one of these from Ebay for 10% of the price next year.

  • source
  • hideshow 19 child comments
  • [–] 24 points 1 month ago (2 children)

    And that's really why they're hoarding them.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 7 points 1 month ago (3 children)

    https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant

    I'm running gwen 3.6 with 131k context window on a 3090, it's fast enough and about as good as pay to play Claude at work.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 4 points 1 month ago* (last edited 1 month ago) (9 children)

    Any 27b Model you can currently recommend for a 16gb AMD ? Mostly coding tasks but not exclusively.

  • source
  • parent
  • hideshow 9 child comments
  • [–] 5 points 1 month ago (4 children)

    9060xt 16gb is the most cost effective new GPU, but if you're going used look for a V620 on eBay. It's a 6800xt chip but in server form factor GPU with 32GB vram. Can be a bit of a pain to set up but by far the most cost effective option IMO

  • source
  • parent
  • hideshow 4 child comments
  • [–] 6 points 1 month ago (2 children)

    V620s were a good deal when you could get them for $350, now they’re $700+ and no longer a good deal.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 5 points 1 month ago* (last edited 1 month ago)

    There just aren't any good deals any more. Prices for everything have gone crazy in the last few months. For coding LLMs the cloud services may now be the least worst value, by design, until they hike the prices.

    That said, I still just paid way too much for a used graphics card so I could do many things locally, because I just don't want to give the likes of Sam Altman a single penny.

  • source
  • parent
  • [–] 2 points 1 month ago* (1 child)

    For your hardware, the VRam is not enough to run 27b but, I’d recommend Qwen 3.5 9b for image / text to text.

    And I’m planning to experiment with Qwen 3.8 9b for text to text.

    4_k_m quantization is the sweet spot for performance and ram usage.

    Also, I find Llama cpp is better than Ollama in terms of performance.

  • source
  • parent
  • hideshow 1 child comment