▲ 973 ▼ Managers (thelemmy.club) submitted 4 months ago* by inari@piefed.zip to c/whitepeopletwitter@sh.itjust.works 180 comments fedilink hide all child comments
[+] AtHeartEngineer@lemmy.world 1 point 4 months ago (3 children) [deleted] permalink fedilink source parent hideshow 3 child comments replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago (2 children) I run 27b at q8 with unquantized KV cache and 256k context on two Instinct MI60 GPUs. Definitely the best model that I have been able to run locally at a reasonable speed. 35b generates tokens as fast as you'd expect from any cloud provider. 27b is slower than 35b, of course, but token generation is still faster than my reading speed and suitable with coding agents. permalink fedilink source parent hideshow 2 child comments replies: [+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[+] AtHeartEngineer@lemmy.world 1 point 4 months ago (1 child) [deleted] permalink fedilink source parent hideshow 1 child comment replies: [–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent
[–] theunknownmuncher@lemmy.world 1 point 4 months ago The wattage is actually relatively low compared to a lot of current gen GPUs (mainly NVIDIA ones). They are software capped to 225W, but the GPUs can handle 300W. Compared to 5090 which is like 600W permalink fedilink source parent