▲ 218 ▼ Apple M7 Ultra Chip Planned With Up to 1.5 TB of Unified Memory (www.techpowerup.com) submitted 2 months ago* by inari@piefed.zip to c/technology@lemmy.world 61 comments fedilink hide all child comments
[–] lepinkainen@lemmy.world 12 points 2 months ago (4 children) Have you tried running a local model on a M series Mac? permalink fedilink source parent hideshow 4 child comments replies: [–] mysteryhumpf@feddit.org 3 points 2 months ago (3 children) Yes ofc I ran Gemma 4 for example, but compare that to the speed of Gemini in the cloud the difference is massive. permalink fedilink source parent hideshow 3 child comments replies: [–] irate944@piefed.social 11 points 2 months ago (2 children) How much RAM do you have and which version of the model did you run? Local LLMs can be just as fast as long your device clears the requirements. If you noticed a huge difference, there’s a really good chance that you tried to use a model that requires more RAM than you have permalink fedilink source parent hideshow 2 child comments replies: [–] mysteryhumpf@feddit.org 4 points 2 months ago (1 child) I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence. permalink fedilink source parent hideshow 1 child comment replies: [–] irate944@piefed.social 5 points 2 months ago That is a bit too much for your hardware, even the Q4_0. You needed a smaller version (26B likely would suit you better. It would be faster and is a MoE) permalink fedilink source parent
[–] mysteryhumpf@feddit.org 3 points 2 months ago (3 children) Yes ofc I ran Gemma 4 for example, but compare that to the speed of Gemini in the cloud the difference is massive. permalink fedilink source parent hideshow 3 child comments replies: [–] irate944@piefed.social 11 points 2 months ago (2 children) How much RAM do you have and which version of the model did you run? Local LLMs can be just as fast as long your device clears the requirements. If you noticed a huge difference, there’s a really good chance that you tried to use a model that requires more RAM than you have permalink fedilink source parent hideshow 2 child comments replies: [–] mysteryhumpf@feddit.org 4 points 2 months ago (1 child) I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence. permalink fedilink source parent hideshow 1 child comment replies: [–] irate944@piefed.social 5 points 2 months ago That is a bit too much for your hardware, even the Q4_0. You needed a smaller version (26B likely would suit you better. It would be faster and is a MoE) permalink fedilink source parent
[–] irate944@piefed.social 11 points 2 months ago (2 children) How much RAM do you have and which version of the model did you run? Local LLMs can be just as fast as long your device clears the requirements. If you noticed a huge difference, there’s a really good chance that you tried to use a model that requires more RAM than you have permalink fedilink source parent hideshow 2 child comments replies: [–] mysteryhumpf@feddit.org 4 points 2 months ago (1 child) I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence. permalink fedilink source parent hideshow 1 child comment replies: [–] irate944@piefed.social 5 points 2 months ago That is a bit too much for your hardware, even the Q4_0. You needed a smaller version (26B likely would suit you better. It would be faster and is a MoE) permalink fedilink source parent
[–] mysteryhumpf@feddit.org 4 points 2 months ago (1 child) I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence. permalink fedilink source parent hideshow 1 child comment replies: [–] irate944@piefed.social 5 points 2 months ago That is a bit too much for your hardware, even the Q4_0. You needed a smaller version (26B likely would suit you better. It would be faster and is a MoE) permalink fedilink source parent
[–] irate944@piefed.social 5 points 2 months ago That is a bit too much for your hardware, even the Q4_0. You needed a smaller version (26B likely would suit you better. It would be faster and is a MoE) permalink fedilink source parent