I don’t know why you think this is true.
I mean you can just try it with DeepSeek or any other large model yourself. This is literally a solved problem now.

Absolutely not.
Evidently you need to read up on how reasoning chains work.
Not really. It might have similarities, but I would never say it’s the exact same problem.
It literally is the same problem. Your brains didn't evolve to do formal logic natively. We emulate it exactly the way the LLM does.
But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.
No, for the same fundamental reasons. It's got nothing to do with state being destroyed either. It has to do with the fact that stochastic systems aren't a good fit for doing symbolic logic.
It’s certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM.
Fact based reasoning is something our brains are famously terrible at doing actually. That's why we use tools like computers in the first place. Our brains can be trained to express patterns of formal logic, and an artificial neural network can be trained to do the same thing. That's why modern LLMs can reliably tell you the number of R's in strawberry.
It also requires understanding how things actually work.
It doesn't, that's the whole beauty of genetic algorithms. All you have to do is specify your selection pressures and your goal criteria, and the system evolves a solution to fit the shape your desire. The LLM doesn't need to get better at doing math, the stochastic approach means it converges on a solution given the right environmental pressures. And that's why hallucinations don't matter, they get weeded out by the attempts being tested against the environment.
Which is the same way it gets better at coding and yet still can’t actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it’s eye watering.
I can tell you haven't actually worked with these tools recently.
Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.
As a communist, I expect you to understand the concept of quantity transforming into quality.
Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis.
Do explain how this is different from saying that human brains are fundamentally limited in that neurons are just next state predictors.
Yes. Inference -> fitness check -> iterate. Agentic retry. It’s incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test.
Except it's not incredibly expensive because the system works on the principle of gradient dissent. It isn't just producing a random value each turn, it produces a plausible value within the context which is precisely what allows it to quickly converge on a solution.
Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn’t change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.
Exactly the way the neurons in your brain are stochastic next state predictors.
Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.
I would urge you to spend a bit of time actually learning about the technical reality instead of continuing to argue here.