It's more than just that. Smaller models are improving rapidly and are already "good enough" for the majority of consumer use cases. It is entirely realistic that in a few years, there will be virtually no demand for trillion-parameter LLMs. Demand for inference might not shrink in terms of functionality, but it will absolutely shrink in terms of memory and compute requirements. And then there will be a hell of a lot more server capacity available than anyone needs or wants, because they're over-investing like crazy now.
What happens when hundreds of billions of dollars spent on datacenters basically goes *poof*?