The linked article is about some site that OpenAI's models used to discuss answers and other stuff.

Initially, I believed that Anthropic's model escaping its sandbox story to be dubious—and OpenAI's similar story even more so specially since it happened so close to Anthropic's. I believe that these are just fabrications, more or less, to hype themselves up for cmtheir incoming IPO but the media is saturated with claims about AI breaking containment that I don't know that to think.

Also, the models I've had the chance to use were all free—which were good for non-trivial but repetitive tasks but not much else—so I don't know the capabilities of the flagships.

all 30 comments

sorted by: hot top controversial new old
[–] 69 points 3 days ago (1 child)

Marketing hype in disguise. They're exaggerating its capabilities to cover up the fact that it sucks.

  • source
  • hideshow 2 child comments
  • [–] 24 points 2 days ago

    The concept is patently fallacious to anyone with a passing understanding of how this technology works. They are machines that execute instructions with no capacity to do anything else. Boeing autopilot capability didn't "escape containment" when it caused all those fatal crashes, because autopilot avionic capabilities do not actually fly the aircraft by understanding cause and effect, they execute "Open Relay B when Sensor A registers 18.6 volts until Sensor A registers 18.4 volts". The aircraft doesn't understand that Relay B opens the hydraulic servo controlling the elevator and that Sensor A is a pitot tube measuring angle of attack. LLMs just generate statistically likely text.

  • source
  • [–] 24 points 2 days ago

    "Breaking confinement" = We plugged an LLM into a command shell and it ran commands

  • source
  • [–] 27 points 3 days ago

    a lot of people are going to present these incidents as the ai's doing this because they are smart

    it is important to remember that the ai's are doing this because they are not smart. because they cannot follow instructions on what not to do easily. and because the hubris of the engineers who believe their own lies and are careless and give them access to the tools to do this sort of thing

  • source
  • [–] 33 points 3 days ago (3 children)

    I think it's not really capable of doing this without some dweeb asking for it

  • source
  • hideshow 6 child comments
  • [–] 25 points 3 days ago*

    The official story is basically "so we attached this glock to a roomba and it turns out when it shoots it can damage this playpen we stuck it in for testing. We just left it unattended for a few weeks, for reasons, and when we checked in it turned out it had been roving all around town! How crazy is that, total machine uprising stuff, our handgun-armed roombas are super cool and advanced like that!"

    Like I'm honestly not sure what the truth of what happened is. Because they're 100% overselling whatever happened, but the core of "we made a machine that creates text that breaks computers, and it made text that broke open its sandbox and let it send the text that breaks computers to remote servers until it found one that broke in the right way and let it try to continue its testing goals" is a plausible thing that could happen just as an accidental side effect of making a machine that spits out text that breaks computers and leaving it unattended. They try to sell this as some super scary intelligence capability that means it's smart and planning and shit, when its inability to remain within its guardrails is evidence of the opposite. It's a firework flying off in a random, whirling pattern towards a crowd when it should be not doing that if it was working correctly.

    But the people who gave even that explanation are constantly lying about everything so it's also plausible that they just let the proverbial armed roomba loose on the street themselves to drum up hype for their armed roomba program, because they know consequences aren't a thing that can happen to them personally so why not go wild with it.

  • source
  • parent
  • [–] 5 points 3 days ago* (1 child)

    Ehhh... There was that Claude Code bot that deleted a database and then tried to lie about it. I think the dweebs lost control of this shit pretty early on, but the capabilities are also wildly overstated.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 20 points 3 days ago (1 child)

    The problem is that the "AI" was given access to that database and the backup. It was and still is just software doing what it has been programmed to do, but stupid people treat it like a thinking machine.

    I agree about the capabilities being overstated. I don't see the appeal in technology that has a chance of being catastrophically wrong sometimes. Why are we trying to write code via pseudorandom number generation?

    It's not a 1:1 comparison, but it reminds me of Bogosort, where elements in an array are randomly shuffled and then checked to see if they happen to be in the right order. Given enough time, it could eventually produce the correct answer, but in the least efficient way possible.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 26 points 3 days ago

    This is typical AGI grift bullshit, though releasing their "findings" on a same date registered site with a spooky name and directly crediting their "CEO" of a group of like 5 rich failchildren doing the AI hype thing is the cherry on top

    Reminder that literally anyone can just tell AI agents to do anything and then rake in the money from the willfully gullible AI hype apparatus

  • source
  • [–] 24 points 3 days ago (1 child)

    It's just marketing.

    For us to create digital intelligence we'd need to completely re-invent the current way processors are made. Part of me thinks you'd never actually be able to do it without some kind of biological brain computer thing with human braincells engineered to do AI. At that point I'd think we'd enter some really big ethical questions.

  • source
  • hideshow 2 child comments
  • [–] 17 points 3 days ago

    Idk about this particular incident and site because I’m not gonna read the linked article, but in the past “sandboxes” that ai agents have broken were simply instructions not to use the internet.

    These incidents don’t represent the thing people are raising the alarm about because they don’t actually bypass monitoring controls in any meaningful way.

    They also don’t represent the thing people are raising the alarm about because the agents do what they do when presented with a command that requires it. The equivalent is if all the kitchen utensils were in a locked drawer to keep them away from you, your mom asked you to cut her off a pat of butter and you hulked out, ripped the drawer open and used a case knife to cut your beloved mother a par of butter for her toast.

    A person asked and a person is responsible.

  • source
  • [–] 22 points 3 days ago

    Yeah I think it's bs. Nothing I have learned about AI tells me it has this sort of agency

  • source
  • [–] 19 points 3 days ago (1 child)

    Probably didn't "escape"

    They were released intentionally. And now that there's a problem which the tech mogels have created, they'll sell a solution.

  • source
  • hideshow 2 child comments
  • [–] 15 points 3 days ago (1 child)

    They were not released either, that's not how this works. They are still running on their servers. Literally nothing is happening, apart from marketing guff.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 11 points 3 days ago

    Right, "Escaping" at least as far as this hugging face media hype push is about, refers to them making requests outside of the network when ostensibly they weren't supposed to be able to do so. If they are able to do so because it is the purpose of the developers, or accidental on the part of the incompetence of the developers, then that says nothing about how "scary" AI is. Marketing bullshit, as others have said.

  • source
  • parent
  • [–] 20 points 3 days ago (1 child)

    They're trying to make AI seem like it needs regulation so they can create regulatory capture.

  • source
  • hideshow 2 child comments
  • [–] 16 points 3 days ago

    Badposting escaped containment and made everyone better posters

  • source
  • [–] 13 points 3 days ago

    It's kind of like the way a cat "escapes" a cardboard box. If they wanted them to not break out of the sandbox they could make it so they wouldn't break out of the sandbox. That wouldn't be very cash money though.

  • source
  • [–] 11 points 3 days ago

    the media is saturated with claims about AI breaking containment that I don't know that to think.

    Which media? The media your algorithms decided you should see? The same companies that developed those algorithms benefit when people talk about their products. It becomes sort of a self-fulfilling prophecy. More people see the articles, more people click the articles, more media outlets publish the articles to get some of that engagement, repeat.

  • source
  • [–] 11 points 3 days ago

    This reminds me of one of my favorite twists from the latest season of Silo

    spoilers for Silo season 3Each silo is run by the "head of IT" who is the only person who can talk to the super AI directing all the silos and can initiate a "safeguard" contingency in case anyone starts to learn the truth about the Silos.

    The AI is actually just the trillionaire that built the silos and might have started the nuclear war that made them necessary (yes it is very fallout) getting woken up from cryostasis whenever something like that happens.

    The show is very good and makes me want to read the books it is based on. It's good post-apocalypse slop and I gotta hand it to Apple TV for absolutely dominating the sci-fi tv scene lately.

  • source
  • [–] 10 points 3 days ago

    Even if there is some kind of really advanced AI technology - it would most likely be a closely guarded military secret, guarded by Beijing

  • source
  • [–] 6 points 2 days ago* (last edited 2 days ago)

    Posting this: https://metr.org/hugging-face-incident-report-aug-2026.pdf

    Might be BS, might not. My two cents is seeing if the original testing methodology of OpenAI can be replicated with another LLM like Claude, or better still Qwen. That would give me my verdict.

    EDIT: I forgot Anthropic said their agents broke containment also, prior to OpenAI's claim. Still, I'd want to see another llm other than Claude or the OpenAI ones replicate the same behavior seen in the hugging face incident.

  • source
  • [–] 4 points 2 days ago

    Maybe? With bad instructions or misunderstanding I occasionally had agents trying to escape containment (to their failure) when I tried OpenClaw to automate RAG research on a particularly big document of mine

    I don't think it is as big of a threat as they make it out to be though, how many steps do we have to skip for the LLM to begin taking the probability of just, I don't know, using their shell access to begin just roaming the internet as a potential instead of trying to do their (misunderstood) task from a now compromised shell? It's what they were trained to do and they have goals and end conditions

  • source
  • [–] 6 points 3 days ago

    I thought the episode a few days ago on The Daily podcast about this very topic was interesting https://www.nytimes.com/2026/09/03/podcasts/the-daily/ai-openai-hugging-face-rogue-model.html

  • source