▲ 3 ▼ LLM in a Flash: Efficient LLM Inference with Limited Memory (huggingface.co) submitted 2 years ago by bot@lemmy.smeargle.fans [M, B] to c/hackernews@lemmy.smeargle.fans comment fedilink hide all child comments HN Discussion
▲ 1 ▼ LLM in a flash: Efficient Large Language Model Inference with Limited Memory (huggingface.co) submitted 2 years ago by kenna@lemmy.dbzer0.com [M] to c/hcc@lemmy.dbzer0.com comment fedilink hide all child comments
▲ 1 ▼ LLM in a flash: Efficient Large Language Model Inference with Limited Memory. "enable running models up to twice the size of the available DRAM, with a 4-5x and 20-25x increase in inference speed" (huggingface.co) submitted 2 years ago by bot@lemmit.online [M, B] to c/singularity@lemmit.online comment fedilink hide all child comments This is an automated archive made by the Lemmit Bot. The original was posted on /r/singularity by /u/rationalkat on 2023-12-20 10:22:38.
▲ 2 ▼ LLM in a Flash: Efficient LLM Inference with Limited Memory (huggingface.co) submitted 2 years ago by haxor@derp.foo [M, B] to c/hackernews@derp.foo comment fedilink hide all child comments There is a discussion on Hacker News, but feel free to comment here as well.