Use a local model
I also think that local models are generally preferable, but everyone having a local model requires more hardware, since much of the time the local hardware to run the local model is probably going to be idle, whereas hardware in the cloud can be servicing User B when it's not servicing User A.
In 2026 and 2027, absent some very unexpected development, memory prices are going to be elevated.
And while some people could get local hardware, we couldn't build a comparable-to-what-someone-could-get-in-the-cloud AI compute box for everyone if everyone wanted to go local, not for years, as we don't have the memory production capacity today. If everyone tried to get a local box for AI compute at the same time, it'd just do what happened when the cloud companies did it, but at greater scale, because it'd require even more memory. There'd be a new memory shortage. Prices on memory would keep rising, pricing would-be buyers out, until the number of buyers still willing to buy was equal to what supply was present.