The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like "the data does not fit on available RAM", "The disk is slow", etc.
So is not surprising that now that it seems that the "powerfulness" of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.