Scientists Train AI to Be Evil, Find They Can't Reverse It::How hard would it be to train an AI model to be secretly evil? As it turns out, according to Anthropic researchers, not very.

you are viewing a single comment's thread
view the rest of the comments
[–] 26 points 2 years ago* (last edited 2 years ago)

Seems like a weird definition of “evil”. “Selectively inconsistent” might be more accurate.

  • source