It is not a model of objects. It's a model of words. It doesn't know what those words themselves mean or what they refer to; it doesn't know how they relate together, except that some words are more likely to follow other words. (It doesn't even know what an object is!)
When we say "cat," we think of a cat. If we then talk about a cat, it's because we love cats, or hate them, or want to communicate something about them.
When an LLM says "cat," it has done so because a tokenization process selected it from a chain of word weights.
That's the difference. It doesn't think or reason or feel at all, and that does actually matter.