From what I understand the breaking containment thing is just hype. The summaries I have seen from people that know more than me, they were testing the models by having them do hacking challenges, and one strategy for second/third place in hacking challenges is to hack the first place team rather than the target to get the requisite data for the challenge. So that was a strategy in the training data, and since the company didn't anticipate that and didn't put as much effort into securing the competing models from each other, that worked. So that is the origin of this "Breaking containment and communicating with the other AIs" claim that is being hyped.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: