So.. not to go too far down the rabbit hole on this topic.. my question is, how does the hallucination occur?
Is it a function of the language syntax within the data set or the contextual meanings of the words being lost as the LLM is compiling whatever answer it's been asked to give?
Just from looking at some of the more humorous screengrabs of ai answers that get posted here.. I can see that a lot of it appears to be poisoned data scraped from shit sources..
However, I have also seen a video of a guy trying to get the ai he was talking to, to tell him how many letter r's appear in the word "strawberry" and it could not get it right.. even as the guy was walking the ai through the spelling of the word and literally counting aloud each letter r as the machine spelled it out.
How does that breakdown happen in a machine model that was probably fed an Oxford dictionary of the English language before anything else?