you are viewing a single comment's thread
view the rest of the comments
[–] 21 points 3 days ago* (1 child)

It's not magic these models are big and you can't just squeeze them onto smaller devices. Quantisation isn't magic either it's literally reducing the accuracy of a larger model to fit on a smaller device reducing its performance. If the data doesn't exist in the model weights then the model will perform worse and you can only compress data so much. The only meaningful tech shift in the space is unified memory but it's not going to be enough for anything this is fundamentally going to need silly quantities of ram for any effective model to run.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 5 points 3 days ago (1 child)

    The main limitation is just RAM on the GPUs. Unfortunately the bubble itself drove up prices on that exact thing dramatically but this not a particularly expensive thing normally.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 9 points 3 days ago (1 child)

    normally.

    GPU prices shot up in ~2018 and 2020 and [etc] and never came back down. Nowadays a 9060XT (the useful version, mind) costs on its own what a decent entry level PC used to. We are never going back to normal.


    Remember to enable JPEG-XL support in your browser, even on your phone!

  • source
  • parent
  • hideshow 2 child comments