you are viewing a single comment's thread
view the rest of the comments
[–] 3 points 9 months ago* (1 child)

You can have applications where wall clock tine time is not all that critical but large model size is valuable, or where a model is very sparse, so does little computation relative to the size of the model, but for the major applications, like today's generative AI chatbots, I think that that's correct.

  • source
  • parent
  • hideshow 1 child comment