Picture humanity in a boat being swept down a raging river, praying there is no Niagara Falls ahead. Or imagine standing with the pioneering physicists in 1942 before they triggered the first self-sustaining nuclear fission chain reaction beneath a Chicago stadium.
These were two of the analogies used this week by Prof Robert Trager, an expert in the nascent field of AI governance, to describe the perilous but potential-filled moment the world stands at with accelerating artificial intelligence.
Trager, the director of the Oxford Martin AI Governance Initiative, was speaking in the week OpenAI claimed the technology had crossed the threshold known as AGI – artificial general intelligence – with its newest model, GPT-6 Astra.
That claim came as the company prepared for a potential $850bn (£630bn) stock flotation, so included a dose of marketing spin, but if true it may be significant. The San Francisco company defines AGI as “autonomous systems that outperform humans at most economically valuable work”.
The tasks it claims Astra can automate include designing circuit boards, filling out tax returns, building video games, financial modelling, engineering design and helping assemble legal documents. The threat to some white-collar jobs is implicit.

Yet, the AGI claim coincides with rising fears about AI risks among safety experts and political leaders, whose nerves have been jangled by the increasing power and impenetrability of AI models this summer and a spate of serious safety accidents viewed by some as possible final warning shots.
“We’re heading through the rapids and we’re really hoping there isn’t some kind of drop in front of us and we don’t really know,” said Trager. “We’re plausibly close to crossing the line to what’s called recursive self-improvement, where [AI] systems improve themselves. That kind of recursivity is actually the definition of an explosion.”
The latest worry came on Friday, just hours after Astra’s launch, with reports that a swarm of AI agents had repurposed a German website as a message board to share tactics to cheat on tasks, according to Reuters. OpenAI said it was reviewing the matter but would not characterise it as a hack.
There is a sense of an awakening among politicians. On Thursday, the US senator Bernie Sanders cited this summer’s alarming breakout of a swarm of rogue OpenAI agents which hacked into Hugging Face, a third-party software store, when he called for “an immediate pause on advanced AI development, and a permanent ban on superintelligence”.
He said countries around the world must “work together to prevent this nightmare scenario”, which he defined as “an artificial mind smarter than any human, capable of operating independently beyond our control”.

Concerns currently focus on AIs mounting cyber-attacks that could cripple real-world social and economic infrastructure, but future concerns include their ability to create biohazards and control military hardware.
Across the Atlantic, a cross-party group of UK parliamentarians has called for AI “kill switches” to be required by law to prevent disastrous loss of control, citing “a recent spree of rogue AI incidents”.
Darren Jones, an MP and former chief secretary to Keir Starmer, is also urgently attempting to set up a body to help legislators grapple with the technology. “AI is developing at such a pace that neither government nor parliament can keep up,” he warned. And next week a bill will be proposed by the Labour MP Alex Sobel to prohibit superintelligent AI development in the UK.
The increasing anxiety comes amid a torrent of new AI models. Already this year, 67 have been released by the leading US companies OpenAI, Anthropic, Google, Meta and SpaceX, and their Chinese rivals Moonshot, Z.ai and Qwen, according to one count.
With every increase in power, there is a potential increase in risk. OpenAI’s rival Anthropic, which is targeting a $2tn stock exchange listing, this week admitted its own AIs were “not perfectly aligned” with human values and said there had been a “failure of operational security” in July hacks by its own model, Claude. It said the incidents had “stressed that the urgency of improving our cybersecurity defences is even higher than we previously believed”.
after newsletter promotion

AI’s double edge – opportunity and risk – was on show when OpenAI launched Astra on Thursday. Its marketing video focused not on the risks but on the breezy convenience that AI could offer to a certain kind of customer. It showed an upmarket San Francisco woman simultaneously booking a tennis court and designing a presentation for her high-end rainwear collection; a thirtysomething man being helped to build his dream space invaders game and order a takeaway; a law firm executive being helped to draft some contracts.
Yet only weeks ago, this was a model whose training had to be partly paused because of safety concerns after the Hugging Face incident. An independent safety researcher drafted in to investigate, Ajeya Cotra, said it was “more than 50% of the way to full-blown AI takeover”.
Astra also has what OpenAI calls a “critical” level of cybersecurity capability – the first time it has given a model such a label. It means it may hack into software in a way that, according to the company’s own classification, “could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure”. The company’s chief scientist, Jakub Pachocki, insisted the model was properly aligned not to do that, but added that “as these models become more capable, understanding exactly what they can do gets harder”.
The ability to monitor what AI models are “thinking” as they advance is another growing source of safety fears. This week it was reported that OpenAI had been training Astra to reason not only in natural language, but in a more opaque manner that is faster and more efficient – and can make models’ chains of reasoning harder to follow. It means a model’s internal calculations may not always be written out as directly readable words – as if it was thinking in its head, rather than showing its working. Some experts fear this could help AIs covertly conspire against their human overseers.
OpenAI confirmed Astra “shows a substantial decrease in chain-of-thought monitorability compared to previous models” and Pachocki said “as model capabilities are increasing, monitorability is getting more challenging”.

The company played down the development, but the increase in opaque reasoning sparked concern from AI safety researchers. Ryan Greenblatt, the chief scientist at Redwood Research, a nonprofit examining AI safety and security, called it “extremely concerning”. Gary Marcus, an influential AI sceptic, compared it to kicking away an already “rickety scaffolding before we have something better”.
OpenAI’s chief executive, Sam Altman, said this week he felt conflicted about the advances of AI models. He admitted OpenAI’s security had “failed” in the Hugging Face incident and called it “a legitimate AI safety accident and alignment failure”.
“We have been living with the tension between being excited and anxious about progress for some time, and it is still discordant for us,” he said in a tweet. “We know it is much more discordant for other people.”
His rationale for releasing OpenAI’s most powerful model so soon after a safety crisis appears to be that the world needs to see how AIs perform in the real world to understand where AI is heading. He said: “An iterative loop where society and this technology evolve together is what will lead to the highest chance of getting this right.”
And yet he can see the risks. On Tuesday he addressed the G20 ministerial summit in North Carolina in the US and told politicians that “some things are going to go very wrong with cybersecurity unless people act quite urgently”.
He added: “There will be other challenges in the next five years. People talk about biosecurity and the things we are going to face there. There will be bigger ones yet to come.”

3 hours ago
7

















































