I'm so sick of the humanising language used in these articles.
An LLM doesn't "decide to cheat".
If you instruct a model to try everything, then inevitably, after n million iterations it will do something you didn't expect.
If you train a model to copy code bases and augment them, then when you instruct that model to develop a bot to do whatever thing, is it really that fucking surprising when it does exactly what it has been created for?