What you're describing for coding, though, is alignment between prompts+context and known good solutions. Yes, it's entirely possible to have the LLM produce novel code solutions, just like it can produce novel sentences - stochastically - but that doesn't mean it's getting a greater basis in reality. It means that it is mapping the highly variable request and existing code to it's training corpus and it keeps adding more maps between prompts and valid code solutions via rewards-based training. Which is still a next token generator no matter how you slice it.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: