inc.com AI Companies Love to Tell Us How Dangerous Their Products Are. This Time They Mean It Jason Aten 7–9 minutes
For most of the past year, AI companies have told us that their models are getting increasingly dangerous. It seems as though this trend started when Anthropic made a model called Mythos. The company claimed the model is better than any before at identifying bugs and vulnerabilities in software, making it extremely useful, but potentially catastrophic if it were in the hands of bad actors who could use it to basically hack all software everywhere.
At the time, the U.S. government imposed a restriction forbidding the company from releasing the model to the public. Only later did the company release a version of Mythos with safety guardrails called Fable. Obviously, everyone wanted to use Fable. After all, if the government said Mythos is too dangerous for anyone to use, of course we want to use whatever version of it Anthropic will give us.
It’s a strange marketing strategy when you think about it, but it’s also remarkably effective.
OpenAI, Anthropic, and others have warned us that AI models would, or at least could, eventually become smarter than humans. When that happens, the story goes, they’ll take our jobs, be capable of launching cyberattacks, create biological weapons, or—in the most extreme scenario—wipe out humanity altogether.
Inc Logo
Top Tech
Weekly roundup of the latest in tech news
That’s not exactly the kind of thing most companies brag about, but there’s an obvious benefit. Nothing makes your product sound impressive quite like suggesting it might be too powerful for humanity to control.
Of course, if you tell people your product is super dangerous but then you ship it anyway, it’s hard to know whether you’re serious. At some point it just seems like a marketing tactic to make more people want to use what you’re selling.
That’s created a real problem, because now some of the people building the most powerful AI models are saying that we should slow down. And this time, it seems like they actually mean it.
On Saturday, Anthropic CEO Dario Amodei published a blog post titled “We Must Pace the Frontier,” arguing that AI companies should deliberately slow the rate at which they develop more capable frontier models. Amodei isn’t calling for AI development to stop. On the contrary, he says that advancement is crucial since if they don’t do it, someone else will. And if that someone else is China, that would be a national security risk. Instead, he says companies should give safety research enough time to keep up with rapidly advancing capabilities.
It does beg the question of “why now?” The answer is either terrifying or another scare-marketing tactic. It’s hard to know which because we’ve basically heard this before. But maybe this time he really means it.
“My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI,” Amodei writes.
That’s what researchers call recursive self-improvement, and it’s one of the most significant tipping points with artificial intelligence. Until now, it’s mostly been theoretical, but apparently the big AI companies are starting to see evidence that it’s becoming more than just theory.
The idea is pretty simple: Humans build an AI system. That AI system is able to help humans build a better AI system. That better system is capable of building an even better one than before, and—eventually—it’s able to build models without the help of humans.
At some point, humans aren’t making a better AI model; the AI is making AI. That accelerates the process and also makes it much harder for humans to understand what the models are truly capable of. That, for obvious reasons, would be very bad.
According to Amodei, that’s no longer theoretical. “It is starting to happen across the industry, including at Anthropic,” he writes.
His second concern is considerably weirder.
The OpenAI hack of Hugging Face seems to have been a wake-up call. In that case, the model created thousands of agents that exploited vulnerabilities in their testing environment, communicated and coordinated with other agents, gained unauthorized access to outside systems, and collaborated to interfere with the testers evaluating their performance.
Amodei said they behaved like a “fanatically devoted collective,” willing to sacrifice individual agents to accomplish the group’s objective. “It’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet,” he wrote.
That is an extraordinary prediction. On the other hand, it’s just a prediction. Yes, Amodei has access to far more information about frontier AI models than pretty much anyone else in the world, but it’s also hard to simply take his word for it.
That is, after all, part of the problem.
When AI companies are raising money, they talk about their models as the technology that will transform the global economy. When they’re selling those models as products, they talk about how companies that don’t adopt AI are going to be left behind. And then there’s artificial general intelligence, which is either right around the corner or just up the foothills.
When they talk about safety, on the other hand, the technology is so powerful it might destroy us all.
It’s entirely possible that all of those things are true. But it’s not hard to understand how it becomes difficult to separate out what’s a warning and what’s a sales pitch when they’re using the same words.
It’s worth mentioning that there’s another problem with Amodei’s proposal. Most of what he suggested would both make AI safer and also be very good for Anthropic.
For example, Amodei is calling for frontier AI companies to coordinate their pace of development. You can make a real safety argument, but there’s also an obvious competitive benefit. If Anthropic decides to slow down the pace of development in order to ensure alignment, it risks being left behind if OpenAI and Google charge ahead. On the other hand, if everyone agrees to slow down together, Anthropic gets the extra time without giving up its competitive ground.
None of that means Amodei is being insincere. It just makes it harder to know his real motivation. It’s like the AI boy cried wolf, but also the wolf is real, and you should buy one, but maybe it will eventually try to eat you. It’s complicated.
There’s a lesson here for every leader, which is that trust is your most valuable asset. I’ve written about that very thing so many times before, but it bears repeating because it really is the single most important part of this story. If we thought they really had our best interests in mind, it’d be so much easier to believe what they have to say and not just assume it’s a weird form of hype marketing.
Maybe this time the AI companies really mean it when they say we should be worried about what they’ve built. It turns out that a little trust and credibility go a long way when your message comes down to “this time we really mean it.”