you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 2 months ago (2 children)

I’m on a MacBook with M2, 32GB ram. Literally just tried:

  • gemma4:12b - very slow, unworkable
  • qwen3:8b - very slow, unworkable
  • qwen2.5-coder:7b - slow but workable. Doesn’t use tools properly in OpenCode.

Well, I guess I’ll try again next year.

For context: my home pc is running gemma4:31b just fine. It’s also a beefy ass desktop, though.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 3 points 2 months ago (1 child)

    Are you running an mlx model? If not, try that. My m4 macbook runs qwen3.6-35b-a3b lightning fast. Has its issues, but fast nonetheless.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 2 months ago (1 child)

    What kind of context length can you get with that, and how much ram?

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 2 months ago (1 child)

    I have a model with 64GB of ram. I've limited context to 16k, in an effort to make it more stable, but tbh - it is rather unreliable no matter what I do. With my setup - mlx_lm and webui, it frequently collapses or loops, no matter the settings. I have done a lot of debugging and have concluded it is probably inherent model behavior.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 2 points 2 months ago*

    That's lame about the looping, but ya I don't think that's a mlx issue, I've had it on my desktop with my nvidia card as well. I also tried fussing with configurations, and I was never sure if it was the models or my settings. I was mainly toying around with LLama based models.

  • source
  • parent