So isn't the solution then to just feed the Claude output through a second tiny LLM with the prompt to slightly rewrite the input? I mean you can probably use a 1b model to do that and that can run on practically anything nowadays
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: