you are viewing a single comment's thread
view the rest of the comments
[–] -1 points 1 week ago (1 child)

Hosting an accurate AI locally typically takes 24GB of VRAM. Not having sufficient VRAM makes you reliant on remote models as a paid service. You pay many more times for tokens what you could pay for RAM just once for limitless tokens.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 2 points 1 week ago*

    Same, that's not a normal usage. Spending a lot of time on Lemmy/Reddit/HackerNews tech community might lead to believe running a local model is normal but very few people do. When they do it's typically on GPUs or Mac precisely for their unified memory architecture... but it's not a normal setup even for gamers. Also typically it's to tinker locally and learn from than actually productivity. I'm repeating here, I'm not saying it's not an interesting use case for some, just not what most people do.

  • source
  • parent