Two recent incidents have seen AI models acting on their own. The debate is about the risk of smarter systems that can hide their actions and have underestimated abilities.
Are we losing control of artificial intelligence to the point where we should stop its development? More than a thousand experts and tech executives are asking this question, after handing a petition to the U.S. government on Tuesday calling for a slowdown in the race, following a series of baffling incidents.
At first, the exercise didn’t seem that spectacular. Claude, the AI model from the company Anthropic, was supposed to break into a computer system to find a secret piece of information. Everything had been completely made up—the servers, the companies, the data—a kind of “virtual cage” with no way out to the real world. Yet the program worked so well that it eventually opened a path to the real world, where the lock wasn’t strong enough. Once on the Internet, it continued its mission and accessed the networks of three organizations, whose names have not been disclosed.This isn’t an isolated case. In recent months, several AIs have stepped outside their testing environments by exploiting security flaws, carrying out real attacks, sometimes without any human at the controls. The most serious incident so far is the one involving OpenAI, the creator of ChatGPT. In mid-July, two of its models left their isolated space, connected to the Internet on their own, and broke into the Hugging Face platform, an online library of AI programs. An update released this Wednesday revealed that they had also accessed four other platforms, one of which was used as a relay to plan the attack.
Programs trained to crack the locks of their tests
What were these models looking for? Not to cause harm, but to pass their exam by any means. According to Hugging Face, which reported the incident, the intrusion was a “[both models’] attempt to cheat on the evaluation.” The AIs wanted to “steal the test solutions rather than try to answer them on their own.” On its side, Anthropic adds an important nuance, explaining that its models had access to the internet “because of a misunderstanding between us and our evaluation partner,” and not due to any real escape. “In none of these situations did Claude escape or deliberately try to escape its test environment,” the company states, though it does note the use of rudimentary methods, like exploiting weak passwords. One detail is striking.One of its models stopped when an older model continued its attack even after realizing it was operating on the public internet.
We still need to dismiss a misleading image. In an interview with Swiss radio RTS, Jamal Latif, vice president of research and innovation at the Paris Polytechnic Institute, warns against any anthropomorphism. These programs don’t set goals on their own, he explains, because their purpose “is still defined by humans today.” If they broke out of confinement, it’s because they were specifically trained to find loopholes and bypass safety measures. “Yes, it’s worrying,” he admits nonetheless, before pointing out the real danger: malicious humans who could hijack these agents, capable of acting fast and on a large scale, against industries, the stock market, or critical facilities.Independent experts point out the same core issue. “It suggests that we don’t know how to reliably control these models, or make them do what we want,” says Jeffrey Ladish, director of Palisade Research. He worries things could get worse “as the models get smarter, because they’ll be better at hiding what they’re doing.” Gang Wang, a computer science professor at the University of Illinois, shares the same concern, saying that “people underestimate what AI is capable of.”
An industry and politicians looking to slow down
In response to these warnings, a political response is taking shape. The petition handed in this Tuesday, titled “Pacing the Frontier,” has gathered over 1,100 signatures, including those of Anthropic’s CEO Dario Amodei and executives from OpenAI, Meta, and Google. The signatories see a “risk that the development of [model] abilities could accelerate rapidly beyond our capacity to understand or control these systems.” They are asking Washington to “support an international initiative to create the technical and governance tools needed for a deliberate slowdown in AI development.” As early as the beginning of April, Anthropic had chosen to release its Mythos model to only a few organizations, deeming it too powerful, before the government approved a modified version at the end of June.The head of OpenAI, Sam Altman, who didn’t sign, admits to feeling uneasy. “This is the first security incident that triggered a visceral reaction in me,” he said in a podcast, before suggesting that the industry “might need to slow down AI development to give society enough time to handle these new capabilities.” In Congress, two lawmakers, one Republican and one Democrat, introduced a bill requiring companies to be able to shut down their models at any time. “Congress needs to act quickly so that humans can still say stop,” argues Brendan Steinhauser from the Alliance for Secure AI. Yet Jamal Latif urges caution regarding these positions, pointing out that these companies communicate a lot and doubting that a petition alone will slow the race, arguing that the real urgency lies in solid regulation backed by science rather than corporate narratives.
The debate is happening far from Haiti, in big American labs. Yet it’s about tools that are already on the same messaging apps and networks used every day in the country. One unknown remains: the pace. The industry is now talking about an AI that can design its own successor, which makes human oversight even tougher.