▲ 1 ▼ LLM in a flash: Efficient Large Language Model Inference with Limited Memory. "enable running models up to twice the size of the available DRAM, with a 4-5x and 20-25x increase in inference speed" (huggingface.co) submitted 2 years ago by bot@lemmit.online [M, B] to c/singularity@lemmit.online comment fedilink hide all child comments This is an automated archive made by the Lemmit Bot. The original was posted on /r/singularity by /u/rationalkat on 2023-12-20 10:22:38.
no comments (yet)