I've actually got a theory around what is going to be the standard in like 10 years. Basically I think everything will be local and the way it will be structured is you will have a single "Supervisor" model that manages a plethora of small specialist models. It will load and unload them as needed from system memory and act as the go-between for the user and the model cluster. That way these can run on really low amounts of memory. Like say you have 10 models all need 10GB of ram. You'd need 100GB to run something like that now. But if you had another model that could take that 10GB, reserve it, and then assign it to the 10 as needed and unload and reload them on demand you'd only need the 10 itself. While keeping the same amount of performance.
I think the reason the US doesn't have a chance in the AI race is because philisophically they just don't function like this. They're all about power. That's why they build datacenters. More compute = more better AI in their mind. But efficiency is the name of the game in anything. If you can do the same thing with 10x fewer resources you'll win every time. It's the same reason Iran just beat the US. Sustainable, efficient, and structural thinking.