The whole thing with R’s in strawberry hasn’t been true for a while now.
I don't know why you think this is true. Transformers literally do not have any concept of algorithms (like the ones humans use to count). You have to add a "skill" - a deterministic implementation of an algorithm - and then you need to use RL to tune the parameters until the transformer outputs the tokens required to pipe the output of the skill to the user instead of producing a normal text output.
Turns out you can use RL to get the model to do basic calculation
Absolutely not. That's just not how transformers work. It's a technical impossibility. Maybe you're talking about non-transformer LLMs, like Mamba, but very few people are using Mamba-based LLMs outside of the research field. When we move beyond transformers, yes, counting is something they can do because they're effectively FSMs.
Notably, this is the exact same problem humans have.
Not really. It might have similarities, but I would never say it's the exact same problem.
The way our brains work is also stochastic
Sure.
and we struggle to do complex math in our heads
But for different reasons. Transformers because they are state destroying. Humans because they have limited and volatile working memory.
But of course, we can reinforce train ourselves to get better at it.
We can reinforce train ourselves to get better at math in so many different ways. We can memorize facts. We have metacognition and can often assess when we do or don't know a fact. We can develop cognitive behavioral algorithms that model physical systems. We can develop cognitive behavioral algorithms that model abstract systems. We can combine all of these together in nested, iterative, and recursive structures. It has only a little bit to do with reinforcing the association between abstract symbolic tokens.
And what we typically do is use an external aid like pen and paper to work through problems, which is basically no different from an LLM harness
It's certainly different because pen and paper are most often used to enhance working memory, but the fact-based reasoning and the algorithmic state machines are encoded in our brain which is impossible for a transformer LLM. It remains to be seen what Mamba-likes are truly capable of, but they certainly are addressing many of the critiques I'm raising.
If you hook up an LLM to REPL in octave, then it can do math quite well all of a sudden. [...] Understanding the limitations of LLMs and how to use them effectively requires moving past reductive thinking
It also requires understanding how things actually work. Hooking up the LLM to a REPL and then iteratively fine tuning it changes the model from outputting a stochastic answer via next-token prediction to outputting a stochastic algorithm via next-token prediction (that will answer the question for you). The LLM did not get better at doing math. It was reweighted to answer questions in the form of "here's a solution that will answer your question for you since I cannot answer you because I have no ability to do math", and within that structure you will STILL get hallucinations and can still fuck the model up by posing sufficiently complex or misdirecting word problems. But the corpus for converting word problems to algorithms is massive, and combined with fine-tuning AND ensemble sampling (which increases your real inference cost in multiples) you're going to produce decent algorithms from math word problems relatively consistently.
Which is the same way it gets better at coding and yet still can't actually solve complex problems in design space, constantly has to use ensemble sampling, and constantly has to be told to re-roll the dice whenever the test fails. And that behavior is so costly under the hood that it's eye watering.
While token generation is the base operation, focusing on that is like trying to understand the brain by looking at individual neuron firings.
Not really. Watching individual neuron firings would be equivalent to watching individual parameter weights and the outputs of each step of the transformer. Token generation is literally the entire functioning of transformers.
What’s actually interesting in both cases are the high level patterns that end up being produced which I’d argue are substrate independent
Yeah, patterns are, by definition, substrate independent. But transformers only maintain high level patterns on a per-token basis. High level patterns can and do emerge from weighted parameter space, and in many surprising ways, but they are fundamentally limited in transformers because transformers are, at base, next-token predictors so even though we get emergent high-level patterns that can, for example, sort lists, we STILL get hallucinations specifically because the high-level patterns are ephemeral on a per-token basis. Reinforcement learning and back propagation can only go so far in fitting the parameter space before it results in contention with other outcomes.
Meanwhile, a combination of an LLM with a harness can be seen as a type of a neurosymbolic system. The neural network generates novel patterns, while the symbolic engine provides the rails for it to function within.
Yes. Inference -> fitness check -> iterate. Agentic retry. It's incredibly expensive precisely because it uses next-token predictors to generate an answer with an already-known fitness algorithm and then just re-runs inference until the answer passes the fitness test. It's an automated human-in-the-loop system where a human says "No, that's not right, try again" and we all know how quickly we run of free tokens when we do that, and we've also all had the experience of the damn thing never getting anywhere near close to the solution after a dozen attempts at reprompting.
Yes, modern transformer harnesses do a TON of work and actually make these parrots useful instead of novelties. But it doesn't change the fact that they are fundamentally statistically weighted parameter-space stochastic next-token predictors, no matter how much you add to them.
Instead of arguing against the technical reality, why not focus on the truth about the harnesses - they add a ton of value and make next-token prediction much more useful in some contexts, especially contexts like producing working code.