you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 14 hours ago (5 children)

Yep, I've been using that same model pretty heavily too, running on ExLlamaV3. It's very impressive for such a little model!

  • source
  • parent
  • hideshow 5 child comments
  • [–] 1 point 14 hours ago (4 children)

    I'm really looking forward to seeing what Qwen4 is like. 3.8 was just 3.6 with additional training, my understanding is that Qwen4 is new from the ground up.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 1 point 14 hours ago* (3 children)

    I hope they make a MoE version. I get like 50 tokens/sec with Qwen3.6-35B-A3B, but can't rely on it quite as much as Qwen3.8-27B.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 1 point 13 hours ago (2 children)

    Sadly they don't seem to have 36B-A3B models on their roadmap any more, they didn't do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 13 hours ago (1 child)

    My Hermes instance periodically checks the overall open-weight LLM landscape, and it recently recommended trying Ornith 1.5 35B-A3B.

    I haven't tried it yet, but on paper, it sounds like it has some potential.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 12 hours ago

    Heh, I downloaded that one just recently, I read that it was good at natural prose and I've been working on a little pet project to make a framework for auto-writing short stories based on a simple premise. Haven't tested it extensively yet though.

  • source
  • parent