Im not super well versed in this, but i dont think you can efficiently distrubute the LLM processing like that. Even if it was possible, it would be very very slow. But I guess speed isnt always needed.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments