you are viewing a single comment's thread
view the rest of the comments
[–] 126 points 2 months ago* (last edited 2 months ago) (21 children)

Basically Apple will be building the perfect computers to run local LLMs.

  • source
  • hideshow 21 child comments
  • [–] 96 points 2 months ago (15 children)

    One can only hope that it totally breaks the AI/LLM at industrial scale, so businesses can run their own AI systems with their own data sets.

    No more of this fucking datacanter horseshit.

  • source
  • parent
  • hideshow 15 child comments
  • [+] -17 points 2 months ago (12 children)

    Local LLMs are cool but also pretty slow compared to cloud. If you have to wait half an hour for your Feature while coding you might still opt for the cloud agent.

  • source
  • parent
  • hideshow 12 child comments
  • [–] 26 points 2 months ago (2 children)

    Yes, they are slower. However, I think that the pricing we’re going to see from the cloud providers might be enough to deter quite a lot of people. At least I hope so:

    The fact that we’re already used to blazing speed generation kinda sucks. Local models are a much more sustainable way of unlocking the benefits of LLMs than giant ecosystem- and community-destroying data centers.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 3 points 2 months ago (1 child)

    I also hope that don’t get me wrong, but as I said: Waiting for the LLM agent to finish coding is currently a bottleneck in software development, they don’t pay high salaries for watching the AI code, they will prefer faster agents even if they are expensive, because they are not only paying the AI Company but also the software engineer overseeing them.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 7 points 2 months ago

    I think that is only going to last as long as the AI providers are willing to operate at a loss. The issue is even with the newer higher price points rolled out this year, they're still losing money. The slower AI machines may be the answer once the REAL profit earning price for the use tokens hits the market. I forsee lots of alternative work going on while the small LLM's are cooking the data. We will have to see once these machines start to roll out, what the use for LLMs will be and how it's applied. I am hopeful.

  • source
  • parent
  • [–] 12 points 2 months ago (4 children)

    Have you tried running a local model on a M series Mac?

  • source
  • parent
  • hideshow 4 child comments
  • [–] 3 points 2 months ago (3 children)

    Yes ofc I ran Gemma 4 for example, but compare that to the speed of Gemini in the cloud the difference is massive.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 11 points 2 months ago (2 children)

    How much RAM do you have and which version of the model did you run?

    Local LLMs can be just as fast as long your device clears the requirements. If you noticed a huge difference, there’s a really good chance that you tried to use a model that requires more RAM than you have

  • source
  • parent
  • hideshow 2 child comments
  • [–] 4 points 2 months ago (1 child)

    I ran Gemma 4 31 B quantized so it fits in my RAM. The decoding speed was decent, but if you look at the newest models for example Gemini flash 3.5 they have a decoding speed of 280 token per second, they generate an entire page before my Mac locally generates a sentence.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 3 points 2 months ago* (2 children)

    Yeah. But they're slow because most of us are GPU peasants. If someone were willing to drop $3-5K on a rig, they could probably run decent, dense models at greater than cloud speeds. Hell, with enough black magic, they could do it with less, but they'd have to go deep into the weeds.

    OTOH, $3-5K buys you a shit ton on Open Router, Claude, Chat, Lumo etc.

    The game is entirely rigged for "you will own nothing and be happy about it".

  • source
  • parent
  • hideshow 2 child comments
  • [–] 1 point 2 months ago (1 child)

    If I buy a rig I might as well host it on the internet to get back some of the investment and sell the compute to others … wait a minute that’s cloud!

    It’s always cheaper to have the same hardware serve multiple people than just one.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 2 months ago

    Yes. People occasionally talk about pooling resources / creating a co-op to buy something like this. Easier to split $200K purchase (and probably $10k/month electricity costs) if 200-2000 people chip in. But then...that's just the cloud with extra steps.

    Can't says I've ever seen a co-op like that work but ICBW

  • source
  • parent
  • [–] 4 points 2 months ago* (last edited 2 months ago) (3 children)

    I guess it depends on your definition of perfect - cheap, good or fast.

    This thing is probably going to cost at least $20K USD.

    Edit:

    "Next year's base M7 processor is expected to arrive in the first half of 2027 and will also upgrade memory bandwidth to about 240 GB/s."

    That's...really fucking slow. What's the goal here - CGI, engineering sims, game dev etc? 1.5TB is cool but at 240GB/s that will crawl for AI use.

    Comparison: this is about $100K, for 7.2TB/s, 252GB VRAM (+500GB system ram, so closer to 750GB total)

    https://www.nvidia.com/en-us/products/workstations/dgx-station/

  • source
  • parent
  • hideshow 3 child comments
  • [–] 3 points 2 months ago

    It’s likely wrong reporting, the m5 ultra gets 614 GB/s today. The m3 pro gets over 800. The rumored m5 pro is 1.2TB/s. Likely this will be 2.4 TB/s, but half the reason for high bandwidth in nvidia chips is somewhat offset in the unified scheme. Mlx needs a lot less copy data around when the gpu can just read it directly.

  • source
  • parent
  • [–] 2 points 2 months ago*

    Something got reported on incorrectly because the M5 Max's have 614 GB/s today, and the Ultra M4 machine's (not laptops) are 819 and that's 3 generations behind a M7 if they make an M7 ultra machine.

    $ for $ you'll get more video ram than paying for a 5090, but it won't be as fast and can't train well.

    Before the ram price decable, you could get a 192gb M4 Ultra for ~10k CAD.

  • source
  • parent
  • [–] 1 point 2 months ago

    Only if they can figure out how to emulate CUDA. Or maybe more containers will start providing alternatives if the machines are popular enough?

    There’s still a lot of image and OCR workflows I have to chug through the CPU and it takes forever.

  • source
  • parent