The key point, which is entirely correct, is that next token predictor label is technically true in a narrow sense but it completely misses what LLMs actually do. Pre training is about matching existing sequences, but RLVR means that the model generates its own sequences and learns from the outcomes instead of just imitating the training data.
The chess analogy has nothing to do with chess having finite states. What the article is saying is that a system trained only on a specific set of games would be a next move predictor, but one that chooses moves dynamically based on winning probability uses live context for its heuristic. So the mechanism encoded in that loop is able to discover and reinforce new patterns that were never part of the original training.