I run local models on a macbook pro incidentally. A 32bln param model can do a lot of useful stuff I find. The progress on making the models smaller and faster has been very rapid, and I fully expect that we'll get to a point where you'd be able to run the equivalent of current frontier models on a local machine within a few years. On top of that, we see things like ASIC chips being developed that implement the model in hardware. These could become similar to GPU chips you just plug in your computer.
The tech industry has gone through many cycles of going from mainframe to personal computer over the years. As new tech appears, it requires a huge amount of computing power to run initially. But over time people figure out how to optimize it, hardware matures, and it becomes possible to run this stuff locally. I don't see why this tech should be any different.