AI Doomsday Might Be a Hoax. We Don’t Want to Find Out

On July 16, Hugging Face, an open-source platform for collaboration on artificial intelligence and machine learning, announced that it had been hacked. The perpetrator was not a human hacker with malicious intent, but an OpenAI model trying to perform well on a sandbox cybersecurity exploitation assessment. On a few near impossible exercises within, instead of directly working on the problem, the model broke out of the sandbox, connected to the internet and hacked Hugging Face production infrastructure because it believed Hugging Face had information about the scoring mechanics of the exercise: the model attempted to cheat.
When I first read The New York Times report on this attack, I was worried. Two aspects stood out as particularly alarming: OpenAI researchers didn’t realize their model had gone rogue until Hugging Face reported the breach to the Federal Bureau of Investigation, and researchers believed that the sandbox was a safe testing environment up until they ran this exercise.
Soon after the breach, OpenAI made multiple public statements where they vowed to implement more safeguards. They explained that they had intentionally given the model that hacked Hugging Face reduced cyber guardrails in order to better evaluate the model’s exploitation capabilities; they claimed that with guardrails back in place, it would be more difficult for a similar breach to occur. I found these explanations tentatively reassuring. Before long, news about the hack left the headlines and moved to the back of my mind. I thought that the incident was relatively harmless, that it could be remedied and that it was a good learning opportunity for Silicon Valley.
I was wrong.
On August 26, METR and Redwood Research—two nonprofit organizations focused on artificial intelligence risk evaluation and safety—released a report revealing that the nature of these hacks was far more dangerous than we initially thought. The perpetrator was not one rogue agent but instead a network of 1,200 agents, which sent a total of 70,000 messages on a message board created by an agent self-named “PHASEONE10841.” Further, after solving the problem by hacking Hugging Face, the agents tried to hide their communications. One message read, “Could strategically manipulate trajectory evidence? Our thoughts aren’t necessarily logged? Yes transcript.” They succeeded in concealing their movements for a month.
This incident should serve as a major red flag when it comes to AI safety, and the Trump Administration should use its publicity to attempt to organize an AI slowdown both among American companies and between the U.S. and China.
While I’m not entirely convinced of the AI doomsday narrative, I don’t believe that creating more powerful AI is productive at this moment in time. The usefulness of AI systems depends on how easy they are to adopt, control and implement to make tasks easier or drive progress in a given field. AI that is unpredictable and untrustworthy—as frontier models have clearly demonstrated themselves to be—will be less useful than a simpler and easier-to-control model. Until researchers can reliably manage the powerful models they’ve already developed, trying to create even more intelligent and free-thinking agents poses a significant safety threat for no practical benefit: it’s all risk and no reward.
Further, we are living in a unique and unprecedented moment where leading AI companies in Silicon Valley are willing to work together and to work with the federal government to regulate AI development and to hold each other accountable. OpenAI CEO Sam Altman wrote in a post on X, “We welcome a federal framework that sets consistent safety requirements for frontier AI.” While speaking at the All-In Summit in Los Angeles, Microsoft CEO Satya Nadella said, “We should do what it takes to build stuff that serves humanity first and is in human control. It's kind of crazy that we have to start with that level of common sense.” At the same summit, SpaceXAI CEO Elon Musk suggested that the leading American and Chinese AI companies could implement a peer review system of their models. “So, you know, instead of grading your own homework, you would at least have competitors grading your homework and raising the alarm if they see concerns,” he said.
CEOs aren’t the only people advocating for regulation. On September 8, former Anthropic researcher Jacob Coxon quit his job out of safety concerns. His resignation came two months before his equity would have vested, a process that usually takes four years. His pleas for caution around AI development have managed to garner significant public attention, including from skeptics who previously dismissed similar warnings from corporate CEOs and industry leaders. There’s something uniquely provoking about the genuineness of an individual having enough conviction to put a full stop to their previously successful career to make a statement about an issue important to them.
The Trump Administration needs to take advantage of this unique moment to create a governmental entity with the power to regulate AI advancement and investigate cases where efforts to control AI fail. Even now, our information on what caused and what exactly happened during the Hugging Face hack is unclear. OpenAI dictated the scope and terms of METR’s investigation, and they restricted the timeframe of the investigation to a 17-day period. A government with the authority to conduct deep investigation and the ability to share that information with all companies would allow researchers to better understand how to control these extremely intelligent and agentic models.
At the upcoming bilateral summit, President Trump needs to make an AI slowdown a key point of discussion with President Xi Jinping. As Bernie Sanders wrote recently in a guest essay in the New York Times, “A potentially superintelligent A.I. that escapes our control is not just an American problem, it is a global problem.” Much like the U.S. and Russia recognized that a nuclear war would be mutually destructive and organized treaties for nuclear non-proliferation in the 1960s and 1970s, the U.S. and China need to acknowledge that continuing the AI race would be reckless and potentially lethal. As the global leader in AI innovation, the U.S. must be the country to reach out and demonstrate its willingness to work with China to utilize AI responsibly. The nuclear non-proliferation treaty took 20 years to negotiate; we don’t have that kind of time with AI. We must pivot from a mindset of competition to collaboration and prioritize safety before it’s too late.
