you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 2 weeks ago (3 children)

I’m pretty excited about this one, is it supported by llama.cpp main yet?

  • source
  • hideshow 3 child comments
  • [–] [S] 4 points 2 weeks ago (2 children)

    Not on main yet I think (as of ~2:50PM UTC on 2026-08-26) -- there's a link to the PR for it in my other comment though. Unsloth's fork has that integrated (they submitted the PR). I wouldn't be surprised if something lands quickly in main, but this is a new architecture so may take a bit for people to figure out how to get the most out of it -- bunch of discussion about e.g. SSD offloading for the ngrams and stuff like that in the github thread.

  • source
  • parent
  • hideshow 2 child comments