you are viewing a single comment's thread
view the rest of the comments
[–] 6 points 10 months ago* (4 children)

It does work, but it's not really fast. I upgraded to 96gb ddr4 from 32gb a year or so ago, and being able to play with the bigger models was fun, but it's not something I could do anything productive with it was so slow.

  • source
  • parent
  • hideshow 4 child comments
  • [–] 3 points 10 months ago* (1 child)

    You can have applications where wall clock tine time is not all that critical but large model size is valuable, or where a model is very sparse, so does little computation relative to the size of the model, but for the major applications, like today's generative AI chatbots, I think that that's correct.

  • source
  • parent
  • hideshow 1 child comment