one is that the model does prediction based on the context of the current data it's looking at rather than just whatever data it was trained on
That's always true. That's how neural networks work. The model is a statistical transform from input to output. Neural networks work by taking input to produce an output. The "context of the current data" is just the input. "Rather than just whatever data it was trained on" is meaningless.
The second part is the feedback loop
Yes, as I said, there's a fitness algorithm and it back propagates adjustments to parameter weights. But that's all it can do, because the model is just weighted parameters. It can't learn facts, it can only adjust its probabilities.
The model makes a prediction
Or rather it produces a random output based on the input and its parameter weights
that prediction is tested against the environment
By something other than model itself that has knowledge of what "success" is and what "failure" is
the model gets feedback
In the form of adjustments to parameter weights
and it iterates
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That's a theory.
And that's what grounds it in reality addressing the issue of it producing meaningless outputs
No. That's a leap that has no basis. The training data is no less a part of reality as the current prompt context is part of reality. What you're describing is that the output of the probabilistic transformer gets tested against various forms of curated fitness tests. The problem with is that the only thing one can do with the the results of fitness tests is to change the probabilities of the opaque parameter space. So you can create a fitness test for how many "r"s are in "strawberry" but the results of that test can only be expressed by weight changes. And those fine tuning adjustments are applied to an opaque network of weights that also includes the opaque probabilistic representations of the fitness tests for how many "r"s are in "perrywinkle" and how many "b"s are in "strawberry syrup".
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
The system as a whole behaves similarly to a genetic algorithm where the solution evolves through the cycle of trial and error.
Yes, it uses the same concepts as a genetic algorithm but it the representation is still the problem. Genetic algorithms for path finding are great because they have discrete actions and limited scope. Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn't make it any less a probabilistic next-token generator that can't represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
Incidentally, this is true for humans as well.
Yes, but just because algorithms are analogous doesn't mean they are functionally equivalent. Humans also have an opaque neural network that functionally behaves like a statistical model. But we have more subsystems than LLMs do, we have more dimensions to our encoding, and we have greater self-governing and modification abilities. So while the genetic algorithm approach is useful, it doesn't make the LLM become closer to reality, it just automates a portion of the fine tuning curation process and leaves all the existing flaws intact.