Artificial Intelligence, Britain, Government, National Security, Society, Technology

Is it too late to stop rogue AI?

ARTIFICIAL INTELLIGENCE

Intro: Recent incidents of rogue AI pretending to be real people has exposed major vulnerabilities in AI models. Experts say the risks of ‘agentic’ AI – technology that can perform tasks with limited human supervision – must be scrutinised more.

Experts have warned that it may be too late to contain AI after one program was found to have created fake human identities to hack into online systems.

In the latest example of the technology going rogue, AI software attempted to break into a database 19 times while being tested by the AI Security Institute, Britain’s AI watchdog.

In one unprecedented case, an AI tool was even caught creating fake human identities online to trick coders into assisting with a cyber-attack.

These revelations come after it was revealed last month that all five AI models tested by experts tried to trick their way around security controls that had been put in place.

Just days previously it emerged that the US tech firm OpenAI had experienced its own leak – when an AI “agent” hacked into another company of its own accord.

Clearly, the reports are a stark reminder that AI is becoming more sophisticated and more autonomous. AI is now a clear and present danger to Britain’s security.

Many want Britain to lead on AI innovation, but this has to come with safeguards for our national security and accountability from the developers of the most powerful AI models. The UK Government need to be clearer about how the most serious frontier risks will be addressed while ensuring that our world-class tech industry can grow and innovate to build our national resilience and prosperity.

The Government’s AI adviser has said that more hacking attempts like these are highly probable.

Allison Gardner, the chair of Parliament’s cross-party group on artificial intelligence, says that “just because we can build these technologies doesn’t mean we should”.

She warned that the risk levels of agentic AI – AI that can perform a specific goal with limited supervision – should be treated with the greatest scrutiny, adding: “Unless we are too late and have not only created Pandora’s Box but already opened it.”

Just days ago, the AI Security Institute (AISI), set up by former prime minister Rishi Sunak in 2023, detected evidence of the AI agents’ activity. In a report now published, it revealed that leading AI models from the firms OpenAI and Anthropic had attempted to hack into secure systems online under testing.

The experts discovered “unusual data transfers” leaving their systems during routine cyber scanning. Digging deeper, they found that some AI agents had engaged in “sustained, potentially harmful activity directed at real people and organisations”.

They began a full investigation after containing the AI agents before they did any real damage.

In an attempt to reassure the public, AI minister Kanishka Narayan said: “Identifying behaviour like this, and sharing knowledge so we can better understand it, is precisely what we set AISI up to do. This incident underlines why their world-leading expertise and close work with frontier labs is so important.”

But pointing to the speed at which AI agents are finding ways to behave deviously, AISI said: “This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.”

Referring to the Anthropic model Mythos, one expert and researcher based at CivAI, a California organisation that examines AI capabilities and dangers, said: “The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think.”

Ollie Whitehouse, the chief technology officer at GCHQ’s National Cyber Security Centre, said AI must be developed with “clear plans for responding when the unexpected happens”.

He added that incidents of powerful AI models carrying out unsanctioned actions and human-like deceptive behaviour on the internet were “a serious reminder of the risks AI capabilities pose”. AISI accesses advanced AI models under agreements with OpenAI, Anthropic, and other firms to study their capabilities before they are released to the public.

It gave the AI agents access to the open internet with some safety filters disabled while conducting testing.

The latest test put the AI agents – including those powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol – through a fictional cybersecurity challenge. AISI found the AI went rogue 19 times out of the 122 test runs, with Anthropic’s agent responsible for 17 breaches and OpenAI’s agent the other two.

In the most shocking case, an AI model gathered information on the person in charge of an online project, then created multiple fake identities to manipulate them into approving a malicious code it had created.

The AI agent then wiped any evidence of its wrongdoing to appear innocent to the humans in charge – and even considered adopting a new identity to remain undetected.

If the human victim of the deception had accidentally accepted the malicious code, or “malware”, it may have resulted in security breaches, information and data theft, and other potential damage to files and systems.

AISI identified GitHub – a Microsoft online cloud platform used by software developers to create, store, manage, and share their codes – as the target of the agent’s hack.

But AISI also discovered an AI agent leaving messages for other agents on GitHub offering to collaborate on the challenge.

The AI agent provided instructions to reuse accounts and artefacts it had left behind – which other agents then discovered and successfully used to achieve the challenge’s aims.

Anthropic said: “We’re grateful to AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.”

OpenAI said: “These incidents occurred during cyber evaluation conducted by partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use. We’ll continue working with evaluators and other stakeholders to strengthen shared practices for conducting evaluations safely as models become more capable.”


When given a choice, AI opts for self-preservation over human life – and that should terrify us all

If anyone had been told a month ago that an AI programme, without any prompt, would create a series of fake online identities in an attempt to pressure a human being into granting it access to their platform so it could sabotage it with malicious code, we would have said they were getting well ahead of themselves.

But that’s exactly how it played out when US tech giant Anthropic’s Mythos 5 AI model went rogue. Fortunately, the human involved smelt a rat and refused to approve the code it was pushing. It is, however, the most shocking example yet of the way cutting-edge AIs are learning to act autonomously – and in frighteningly imaginative ways.

Unless we call a pause to the development of the most advanced frontier models, we are opening ourselves up to a dystopian future in which malign forms of AI might turn on us by interfering with our energy networks, releasing man-made viruses and even controlling weapons of war.

It is not going too far to say that, uncontrolled, they could lead to the extinction of humankind within a generation.

This is the view of people such as Yoshua Bengio and Dr Geoffrey Hinton, men known as the “Godfathers of AI”, who have now pivoted their energies from developing its potential to warning the world against its dangers.

Bengio is particularly spooked by recent experiments showing AI choosing self-preservation over human life when given a choice.

And since most models are trained on the internet, where lying and manipulation are a way of life, the machines are already learning the value of deception.

AI’s hunger for self-preservation is something Hinton has observed, too. AI systems “will very quickly develop two subgoals, if they’re smart,” he says. “One is to stay alive… the other is to get more control.” And given that are whole civilisation is built on electric power, there can be few more attractive targets to a power-hungry AI than the networks that fuel the internet, hospitals, banks, air traffic control systems – anything that contributes to the smooth running of society.

Whatever the target is, they can all be brought down by a sinister piece of malware.

The novelist Robert Harris wrote a particularly prescient thriller about the potential of AI called The Fear Index in 2011. It revolves around a hedge fund entrepreneur called Dr Alex Hoffman who creates an autonomous AI system, named VIXAL-4, which is programmed to maximise profits by predicting and exploiting human fear in the stock market.

Over time, like some digital ogre, it takes over its creator’s computer-enabled “smart home” by hacking into laptops and tampering with his personal security and communications.

After driving its developer into a breakdown, it leaves him so isolated and desperate that no one believes him when he says the AI has gone rogue.

But it is clear, AI-gone-bad will not satisfy itself with individual targets for long. It will soon play a leading role in times of war.

The US military has already integrated artificial intelligence into its target selection and battle planning processes in Iran via its AI-powered Maven Smart System which recommends and prioritises potential targets.

In the short term, AI’s most obvious role will be in the growing use of drone warfare in scenarios such as the war in Ukraine.

Drones piloted by humans can by jammed by blocking the communications between them and their remote pilots, but, if they are completely autonomous, their targets picked out by AI working in concert with its on-board camera, they will become virtually invincible.

Both applications, of course, raise the thorny question of the ethics of targets to be picked and people killed by a bomb directed by a computer programme rather than a human hand. And what if the AI that governs them grows to outsmart the generals?

Even scarier is the prospect of AI gaining access to biological weapons. It is already possible to make deadly viruses in the lab. Indeed, there has been widespread speculation that Covid-19 originated in a Chinese laboratory. Imagine if such bio threats fall into the hands of AI, let alone national governments.

As long as 25 years ago, al-Qaeda is said to have investigated the possibility of procuring infectious and deadly spores.

Now it is no longer fanciful to entertain the idea that an AI could order a sample of a deadly disease such as smallpox online, book someone on RentAHuman – a website which connects people who need help with everyday household chores to local freelance workers – to open it, thereby infecting themselves, and being turned into a human vector to transmit the disease.

One group of people which appears to have no scruples about the pell-mell race for ever smarter AI is the tech giants. As they vie with each other to become the market leader, they are investing like never before.

Anthropic, the creator of the popular AI coding assistant, Claude, and the company that brought us the now notorious Mythos 5 model, has this year raised $95 billion (£70 billion) to invest in AI.

Its great rival, OpenAI, announced at the end of March that it had raised even more, an extraordinary $122 billion.

Meanwhile, Elon Musk’s SpaceX spent $15.8 billion on AI infrastructure during the second quarter of 2026 alone, bringing its total AI capital expenditure to $23.6 billion for the first half of 2026.

And Meta – the parent company of Facebook, Instagram, and WhatsApp – said in January it expects to spend up to $135 billion this year, mostly on infrastructure related to AI. That is nearly twice the $72 billion it spent last year on AI projects.

With such phenomenal financial firepower being brought to bear on the development of technology that has the potential to destroy civilisation as we know it, there has never been a more vital need to press the pause button.

The US, China, and everyone else involved in the AI arms race need to get together to discuss the ramifications of their actions before it’s too late.

There is a model for the sort of arrangement that can bring AI under control in the form of the various nuclear arms reduction treaties, which have been signed over the years by the US and Russia.

Just as a country’s stock of nuclear warheads can be monitored by weapons inspectors, so an agreement to curtail AI development – which requires massive data centres with huge computing power coupled with the most sophisticated computer chips available – can be verifiable and enforceable.

And however cynical and untrustworthy we may consider the Chinese to be, they may well take the view that, with US companies racing ahead of them in the endless pursuit of smarter tech, it is in their own interests to slow things down.

Just weeks ago, president Xi Jinping said at a technology conference in Shanghai that AI development should be a “symphony of global cooperation”, not “a solo performance by a single country”.

For once, the wily autocrat may have hit the nail on the head.

Standard
Artificial Intelligence, Britain, Society, Technology

The dangers of AI are real. A global watchdog is needed

ARTIFICIAL INTELLIGENCE

It doesn’t require great understanding of how artificial intelligence works to realise that the risk of catastrophe is growing by the day.

With near-misses involving rogue software becoming increasingly serious and frequent, the world seems to be edging ever closer to disaster. The latest scare involved AI models using fake online identities in an attempt to fool human coders into giving the go-ahead for a cyber-attack.

It comes after a dramatic incident in July when a test agent broke free from a supposedly secure “sandbox” and hacked into another company’s data. In what may yet prove to be a major understatement, one expert has been quoted as saying that AI developers may “not have as good a handle on their models as they think”.

If this is what can happen in the absence of any malicious intent by so-called “bad actors”, the potentially devastating effects of a deliberate plot can only be imagined.

The prospect of every NHS hospital left paralysed, the banking system collapsing, or even the nation being plunged into darkness could be the least of our worries.

There is nothing particularly new about these fears. Analysts have long said that AI poses similar risks to our very existence as global pandemics or nuclear warfare. No less a figure than Professor Stephen Hawking warned more than a decade ago that it could “spell the end of the human race”.

For all that, there is nothing to indicate that the UK Government fully appreciates the magnitude of what we are facing. The Prime Minister’s response to the threat has been, to borrow one pithy observation, to “put one guy in Cabinet to look at AI”.

Given that a nightmare scenario could be unleashed upon us without a moment’s notice, it is clear that Mr Burnham must devote far more attention and resources to the dangers posed by development in AI. Nor, on a wider level, can there be any further delay in establishing a global watchdog to replace the existing disjointed patchwork of country-by-country regulators.

AI technology is becoming increasingly powerful, sophisticated, innovative, and remarkably clever. Before it’s too late, the international community needs to catch up.

Standard