but RLVR means that the model generates its own sequences and learns from the outcomes instead of just imitating the training data.
It generates it own sequences by predicting the next token based on the training data
And then it updates it's weights based on the outcomes of those predictions.
It's still a probabalistic next token generator.
But RLVR allows humans to define deterministic fitness tests and then continuously run the next token predictor and let the deterministic software use the results to backward propagate adjustments to the weights.
Which all presupposes that the output of an LLM is probabalistic and based statistical weights of model parameters, which returns is back to probabilistic token generation by a statistical model that doesn't represent anything resemble knowledge or fitness. In theory, you could tweak parameters with RLVR for a millennia and the LLM will just thrash between various states of meaninglessness, because the conjecture of LLMs is that meaning is exclusively encoded in the statistical distribution of tokens.