“AI systems are unpredictable and difficult to control— we’ve seen behaviors as varied as obsessions, sycophancy, laziness, deception, blackmail, scheming, ‘cheating’ by hacking software environments, and much more. AI companies certainly want to train AI systems to follow human instructions (perhaps with the exception of dangerous or illegal tasks), but the process of doing so is more an art than a science, more akin to ‘growing’ something than ‘building’ it. We now know that it’s a process where many things can go wrong.”
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: