Does it matter whether it’s part of it or done immediately after? For all intents and purposes it’s the same thing. Like I said, if the input and outputs are the same, what does it matter how the process works?
From the user’s perspective, where one question results in many inference calls, it would look like the LLM learns while it works, assuming such training would be enabled, which they obviously wouldn’t but could do.