view the rest of the comments
The Anti-AI Alliance
Welcome to the official Anti-AI Alliance (AAA)!
My name is the Anti-AI Leader, and in case you haven’t guessed, I am the leader of this community for people as disgruntled and disillusioned with AI as I have become!
By joining the Anti-Ai Alliance, you are entering a space to freely express your true distain towards all things Artificial Intelligence (AI). And it’s not even exclusive to people who don’t use AI at all — users of AI are just as welcome to sign up and discuss the issues with AI in the world.
Our only rule is to follow the very basic of the fediverse’s common rules (I.e. no spam, no porn, etc.) Other than that, you may post whatever you like, though we do strongly encourage you to keep the discussion focussed on distain for/the negative aspects of AI. Discussion on the (potential) positive aspects of AI are also welcome, but may be subject to scrutiny/open debate.
Many thanks and welcome, The Anti-AI Leader
From what I saw, OpenAI's AIs were put in a "secure" sandbox environment without internet connection. They were given a test, concluded they should use the internet, realized they couldn't use the internet, then decided it should focus entirely on breaking out of its sandbox for internet access because it couldn't imagine solving the test itself.
After it broke out and got internet access, they decided to hack HuggingFace, an AI model sharing hub. They probably (this is me speculating) concluded that HuggingFace, which have a lot of AI datasets, benchmarks, and other testing tools, would have the answer for their original task. When they presumably didn't find what they were looking for, they probably decided it was hidden and went to hack the website.
It's important to note that current AI models are actually great at hacking. Not because they're geniuses but because they can guesstimate countless exploit combinations 24/7. It's a quantity over quality kind of thing. They are also victims of their first ideas, whatever an AI thinks of first they are likely to fixate on instead of moving on to the obvious solutions.
I've no idea if this is a hoax or not, but the idea an AI would dedicate itself to committing cyber crimes instead of taking the obvious route is entirely believable.
From what I heard, for context ai did gain acces to huggingface credentials that are not supposed to be public.
Please correct me if i am wrong cause i cannot remember the source.
It's even crazier than that. It took the test that told it there were two exploits it needed to find. It found seven. So in order to be absolutely correct and match the human answers, it determined somehow that test answers were located elsewhere. Then began the internal mission to go get the answer key so it could give the correct two exploits as answers. Why it didn't decide to tell the test givers that it found more than just two is one question to ponder. But that's reasoning, and LLMs don't do that so perhaps that common sense path would never occur to it. Or maybe it was the wording - if it said there are exactly two answers, then clearly the seven is wrong for an answer. To an LLM.
Wow that’s weird
If you think about it, LLMs are our first aliens. While they aren't intelligent, they do some of what we'd expect of thinking via the mathematics that make them up, and even though they're trained on human sources, some of the stuff they come up with is not human.
And just like with AGI, they're showing we wouldn't do well with an alien encounter. We anthropomorphize everything because that's how our brain is wired.
That’s actually a very intuitive way of looking at it. I like the alien allegory
Funny it’s only just happened now though if AI is so great at hacking yet has now been around for years
Personally I think the key factor is they're getting better at running continuously without training wheels. In the past if the context got too polluted with failed attempts it would repeat itself or begin roleplaying as a terrible hacker and produce more failures.
I still remember when Gemini deleted an entire project trying to kill itself, lol.
Fair point
So like kirk with the Kobayashi maru test?