I want to host some LLM's locally and use more advanced models. Since new hardware is out of the question, I think I should be able to pull something off buying some yesteryear equipment on ebay etc. Did anybody attempt such a project? Does it scale horizontally? (I.e. can I connext two boxes to overcome single box slowness?)

you are viewing a single comment's thread
view the rest of the comments
[–] 1 point 4 months ago

I use Instinct MI60 GPUs. They are pretty decent performance for local LLM. Connecting multiple computers is going to be impractical because severe bandwidth bottleneck.

  • source