you are viewing a single comment's thread
view the rest of the comments
[–] 12 points 2 months ago (4 children)

Have you tried running a local model on a M series Mac?

  • source
  • parent
  • hideshow 4 child comments
  • [–] 3 points 2 months ago (3 children)

    Yes ofc I ran Gemma 4 for example, but compare that to the speed of Gemini in the cloud the difference is massive.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 11 points 2 months ago (2 children)

    How much RAM do you have and which version of the model did you run?

    Local LLMs can be just as fast as long your device clears the requirements. If you noticed a huge difference, there’s a really good chance that you tried to use a model that requires more RAM than you have

  • source
  • parent
  • hideshow 2 child comments
  • [–] 4 points 2 months ago (1 child)

    I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence.

  • source
  • parent
  • hideshow 1 child comment