Say you had a version of chat GPT that hadn't been trained on a corpus that included enough information about the number of Rs the word "strawberry" contains, didn't have workarounds for character counting, and has context sharing.
You ask it "How many Rs does Strawberry contain?"
It burns a bunch of tokens, returns "Strawberry has 7 Rs. No wait, that's not right, strawberry has 5 Rs. No wait..." etc.
You say "Strawberry has 3 Rs. How many Rs does Strawberry have?"
It says "Strawberry has 3 Rs."
You ask "How many Rs does this exact string contain: 'Strawrberrrry'?"
A new session using shared context responds "Strawberry has 3 Rs."
You say "No, it has a different number of Rs. I didn't ask about Strawberry this time"
It responds "Sorry about that! You didn't ask about Strawberry, you asked about Strawrberrrry. Strawrberrrry has 7 Rs. No wait, that's not right. Strawrberrrry has 5 Rs...."
LLMs, at their core, use models that have computed lexical & semantic similarity to predict words (really they perform contextualized vector transformations, but let'snot get into that). When they "learn", they just do this more effectively using more text or using the same text more efficiently. Next word prediction is not the same thing as learning concepts, like the concept of numbers, non-numerical things being ascribed numerical values (ironically, since LLMs function by turning words into numerical vectors), or counting.
When it responds "Strawberry has 3 Rs" it isn't because it knows what that means or how to gain that knowledge about other words. It is merely parroting back what you've told it because "Strawberry has 3 Rs" in its shared context has very close lexical similarity to your query "How many Rs does Strawberry have?". It also parses Strawberry as ["straw", "berry"] encoded into numerical values representing their relationship in the corpus' vector space (e.g. {[ .420, -.67, .67], [.420, .69, -.69]}) - so unless it has instructions to further break those tokens into characters, it does not have the capability to count. You can actually do some weird math using these vectors and their relationships (their distance apart in the 3D vector space, angles between vectors, etc actually relate to the semantic content of the tokens), but you lose granularity such as the number of letters in a token when you look at words this way.
Learning is more than just computing lexical similarity. Modern LLMs mimic reasoning by expanding the query with intermediary tokens, basically creating a temporary scratch pad of related words on the same vector, but don't actually reason (this is known as chain of thought). Learning is partly computing lexical similarity, but it is also about extrapolating concepts from facts and inferences, applying those concepts to novel simuli, etc. LLMs really don't do reasoning, which is part of learning.