you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 1 month ago

Yes, that's basically what the article is about. They run the LLM across both GPUs.

But that's a feature of llama.cpp. SLI doesn't really exist any more, and NVlink requires a specific setup, which the 4080 is not part of (the 3090 was the last consumer one, apparently). So you couldn't pool the VRAM.

  • source
  • parent