Fears of AI-induced Armageddon are overdone
Applying basic security principles would go a long way to fixing problems identified by Anthropic and OpenAI, writes Ciaran Martin
Economist
Aug 23rd 2026
On May 19th 1998, seven hackers from a group called L0pht walked into a Senate office building in Washington, DC. Granted pseudonymity, “Mudge”, “Space Rogue”, “Kingpin” and four others came in ill-fitting suits to warn a congressional committee that any one of them could make the internet unusable in 30 minutes. Fourteen years later Leon Panetta, then America’s defence secretary, warned of a “cyber Pearl Harbour”.
Cyber-security has a long history of apocalyptic prophecies which, despite lots of individual hacks and harms, never quite come to pass. Now, another 14 years on, there is a new wave of catastrophising. In April Anthropic, a frontier AI lab, said its latest model, Claude Mythos, was so good at hacking it would not be released publicly. In July its main rival, OpenAI, disclosed that several of its AI agents broke out of a testing environment and into Hugging Face, a library of AI models. There followed similar disclosures from Anthropic, Meta and Britain’s AI Security Institute.
One news outlet, capturing the growing sense of foreboding, called it “the start of a dangerous AI cyber era”. So is it different this time? Or does the uneasy equilibrium between attack and defence still hold? Mostly, and thankfully, it does, for two reasons.
First, the pattern in which scary announcements give way to more measured assessments within weeks keeps repeating. Anthropic’s standout claim for Mythos was that it had found a flaw in widely used software dating as far back as the 1990s. Subsequent analysis, however, showed that the threat was much less potent than headlines suggested. Within weeks, predictions of a tsunami of AI attacks had given way to forecasts of merely a stormy period, and then to talk of AI as an opportunity for security improvement. If AI is good at finding weaknesses, it can also help fix them.
There is likewise less to the supposedly “rogue” agents than meets the eye. They were not going rogue. They were doing what humans had told them to do, with the precocity and indiscipline of talented but unsupervised children. They did no real harm.
The incidents differ in detail, but the common factor is a failure to control the testing environment. This must be improved. Some in the cyber-world expressed astonishment at the weakness of testing controls, noting that security-company tests ending in the same outcomes might lead to lawsuits and even prosecutions. That accountability principle is crucial and goes beyond testing. Agents can act autonomously but are not autonomous: whoever sets one to work owns what it does. If the world lets AI agents run riot without supervision or accountability, cyber-attacks will be among the lesser of the ensuing problems.
The second reason for calm lies in understanding hackers. The reports about Mythos and the agents relate to artificial testing conditions. Using agents for hacking in the real world means bypassing guardrails and incurring extremely heavy computing costs. These constraints matter: several years into the age of AI, evidence that malicious hackers are using advanced AI-hacking techniques remains remarkably scant.
Last year a working paper from MIT Sloan, written with an AI-security vendor, claimed that 80% of ransomware incidents (the most harmful form of cyber-crime) were AI-driven. Marketing departments seized on the finding. Respected publications amplified it. But the paper was withdrawn after critics noticed it had counted, among other things, WannaCry—the attack in 2017 in which misfiring North Korean hackers wreaked havoc in more than 150 countries—as AI-powered.
One of the people who spotted this nonsensical claim was Marcus Hutchins, a British cyber expert credited with stopping WannaCry. Mr Hutchins now despairs of the “freakout over Mythos”, given that the “local water treatment plant runs [on] Windows XP”, a Microsoft product already obsolete when WannaCry struck, and “the protocol that routes internet traffic is secured by everyone just agreeing that hijacking it would be uncool.”
Both points are crucial. Mudge and others warned in 1998 that fundamental parts of the internet’s architecture were wildly insecure. Some still are. The reason nobody has brought down the internet is that there is no reason for any capable hacker other than a total nihilist to do it, and whoever did could expect severe consequences. That remains the uneasy basis of modern digital life.
Mr Hutchins’s point about critical infrastructure is more timely still. As the media gawped at AI-testing failures, in the real world hackers were hitting water facilities in a dozen American states. Whoever they were—some say Iran, but President Donald Trump demurs—they were using crude techniques exploiting basic security failures, not advanced AI. Serious supply disruption was only avoided thanks to stored water, manual valves and observant staff.
And there is the sting: underlying cyber-security is often so weak hackers have no need of expensive new techniques. AI doesn’t yet give most hackers magical new powers. But it can help them prepare, speed and scale up. That could make it likelier that security neglect will come back to haunt organisations.
This is the boring but vital lesson. At Black Hat, a security conference held this month, two OpenAI employees presented on the Hugging Face incident. Their concluding recommendation was that applying basic, long-established security principles would have gone a long way towards containing OpenAI’s over-enthusiastic agents. Depressing, perhaps, but if it takes inflated hype about omnipotent AI bots going on hacking sprees to get the West finally to focus on fixing the cyber-security flaws that have long put water and other vital services at risk, then so be it.
Ciaran Martin is a professor at the Blavatnik School of Government at Oxford University. He ran Britain’s National Cyber Security Centre from 2016 to 2020.