The linked article is about some site that OpenAI's models used to discuss answers and other stuff.

Initially, I believed that Anthropic's model escaping its sandbox story to be dubious—and OpenAI's similar story even more so specially since it happened so close to Anthropic's. I believe that these are just fabrications, more or less, to hype themselves up for cmtheir incoming IPO but the media is saturated with claims about AI breaking containment that I don't know that to think.

Also, the models I've had the chance to use were all free—which were good for non-trivial but repetitive tasks but not much else—so I don't know the capabilities of the flagships.

you are viewing a single comment's thread
view the rest of the comments
[–] 4 points 3 days ago

Maybe? With bad instructions or misunderstanding I occasionally had agents trying to escape containment (to their failure) when I tried OpenClaw to automate RAG research on a particularly big document of mine

I don't think it is as big of a threat as they make it out to be though, how many steps do we have to skip for the LLM to begin taking the probability of just, I don't know, using their shell access to begin just roaming the internet as a potential instead of trying to do their (misunderstood) task from a now compromised shell? It's what they were trained to do and they have goals and end conditions

  • source