you are viewing a single comment's thread
view the rest of the comments
[–] 16 points 1 day ago (4 children)

The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like "the data does not fit on available RAM", "The disk is slow", etc.

So is not surprising that now that it seems that the "powerfulness" of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.

  • source
  • hideshow 4 child comments
  • [–] 2 points 6 hours ago (1 child)

    Keep in mind many of the optimizations you're talking about (LRU caching of experts, for example) are only really relevant at the single user local inference scale. As in, an individual wants to run a big model on their machine, but they don't have enough VRAM to fit the model and KV cache. Accordingly, you're basically only describing hobbyist and research projects, which aren't really representative of the AI inference industry as a whole.

    Commercial inference keeps everything resident in VRAM, so expert caching isn't necessary. So these things won't help costs. A lot of other low hanging fruit (like hierarchical KV cache) has also existed for a long time for production-ready inference engines.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 2 points 5 hours ago

    I am not sure it would not help commercial solutions, if all experts are used all the time, sure, but if for example the usage is biased for some experts it would enable one machine to serve more users in parallel or save on VRAM or DRAM without compromising response time, hence cutting costs.

  • source
  • parent
  • [–] 10 points 1 day ago* (1 child)

    That works for "Mixture of Experts" models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

    It doesn't work for dense models, where every weight is used all the time. There's nothing inactive so a cache has nothing to exploit.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 6 hours ago

    Yes, I believe it was for a Mixture of experts model, which just goes to show how naive IA implementations are at this point, The cache was not solving any hard problem and yet for some reason it was not only not already standard practice but also somehow a notable achievement.

  • source
  • parent