Or rather it produces a random output based on the input and its parameter weights
The bias is precisely what makes it not random, but rather stochastic. There's a very big difference here.
In the form of adjustments to parameter weights
I'm talking about feedback from the environment it operates in. That's the actual test that allows the model to keep adjusting outputs towards a specific target rather than them being random. And that's what makes the whole thing useful in the end.
The training apparatus outside the model does this repeatedly, yes, under the thesis that tweaking parameter weights will result in fewer failures to the deterministic fitness algorithm. That’s a theory.
No, that's not a theory, that is precisely what we measurably observe in practice with coding harnesses. And having built one myself, I can tell you for a fact that this works exactly the same way a genetic algorithm does, and large part of making an effective harness comes from ensuring that the model gets actionable feedback.
No. That’s a leap that has no basis.
The basis is me having worked on a harness and observed how the model outputs improve based on the feedback. There's also plenty of research on the subject explaining how and why this works in detail. The parameter space is also not nearly as opaque as you seem to think.
At no point is the LLM getting closer to learning facts, and the thesis that knowledge or skill is representable as a statistical model is unproven and seems increasingly unlikely.
That's missing the point entirely. The question isn't about whether LLM is getting closer to learning facts. It's about whether the biasing from the feedback loop causes the LLM to produce relevant outputs. Also, the thesis that knowledge or skill is representable as a statistical model is very much demonstrated by world models where a temporally consistent simulation of the environment is maintained.
Applying the same technique to fine tuning an LLM is a better use of time than manually fine tuning, but that doesn’t make it any less a probabilistic next-token generator that can’t represent stable facts and rules and where every fine tune for one input is always in tension with the fine tubes for all other inputs.
That's not how any of this works at all. You're not trying to get it to represent stable facts, you use things like compilers, test harnesses, formals specs, and so on, to create the selection pressure. Then the model is the stochastic part of the system which finds a path that satisfies the selection criteria. Or, with robotics, you have models interact with the physical world and use the feedback to adjust predictions within the model.
Yes, but just because algorithms are analogous doesn’t mean they are functionally equivalent.
Yet, they are functionally equivalent in accomplishing many tasks now. And of course, biological brains have many more subsystems and are more complex in general. I'm not arguing that part at all. My point was that what grounds our mental models in reality is the same feedback loop we use to ground LLMs, and it's effective for the exact same reason. I also don't think LLMs are the pinnacle of AI, they're just one piece of the puzzle, and as I noted earlier, people are already moving towards world models now.
I find world models to be fundamentally more interesting than plain LLMs because if a model encodes the rules of how the physical world works, that provides a foundation for meaningful communication. Humans can talk to each other easily precisely because we all have a shared context which is the environment we live in. And we see how the rate of misunderstanding quickly goes up when we start talking about abstract topic because they can be interpreted in many different ways. So, if models can share the understanding of the physical world with us, it becomes a lot easier to tell them what you want, to correct them, and to have them genuinely understand requirements in a human sense.