Skip to content

The case for caution

Nick Bostrom, one of the world’s foremost philosophers on the consequences of AGI and the author of Superintelligence and Deep Utopia, was asked what comes to mind when he hears the word “future.” His answer?

“Hold on for dear life. Buckle down.”

I’ve come to agree with him. Four things are true at once: AGI is closer than most people think, it will soon be able to improve itself, nobody building it has a reason to slow down, and the main justification for racing doesn’t hold up. Any one of these would be cause for concern. Together, they make caution the only rational position.

It’s closer than it looks

The leaders of the frontier labs, who have direct access to the most advanced systems, have grown more confident about short timelines, not less. In November 2024, Sam Altman said “the rate of progress continues.” By January 2025, he wrote, “we are now confident we know how to build AGI.” That same month, Dario Amodei said he was “more confident than I’ve ever been that we’re close to powerful capabilities… in the next 2–3 years.” Demis Hassabis went from “as soon as 10 years” in late 2024 to “probably three to five years away” by January. Ilya Sutskever started a lab on the premise that “superintelligence is within reach.”

They’re joined by independent voices: Geoffrey Hinton and Yoshua Bengio, two of the most influential AI researchers alive; Ben Buchanan, the Biden administration’s top AI adviser; economists, mathematicians and national security officials; and former frontier-lab employees like Leopold Aschenbrenner and Daniel Kokotajlo. Some experts doubt AGI is imminent. But even if you set aside everyone with a financial stake, enough credible people predict short timelines that their warnings deserve to be taken seriously.

When industry insiders ask for regulation, it usually signals regulatory capture: companies shaping the rules meant to constrain them. Here, the evidence points the other way. Never before have so many experts called for safety measures while standing to make so much from moving fast.

Humans have an uncanny ability to be skeptical of new technology:

Newspaper headline: “The Internet doomed to fail over unfulfilled promises”
Daily Mail, 5 December 2000: “Internet ‘may be just a passing fad as millions give up on it’”

Sometimes the skeptics are right. But sometimes the warning signs become impossible to ignore. I believe we’ve reached that moment.

It will soon improve itself

AI systems are increasingly contributing to their own development. Once they’re as competent as human AI researchers, they will likely start improving themselves at a runaway rate, creating the mother of all feedback loops. Remember this when someone confidently claims a capability is still far off.

AI research is unusually easy to automate. As Leopold Aschenbrenner puts it in Situational Awareness, “we don’t need robotics—we don’t need many things—for AI to automate AI research.” The work is done fully virtually and runs into few real-world bottlenecks: “read ML literature and come up with new questions or ideas, implement experiments to test those ideas, interpret the results, and repeat.” And the people best placed to train models for that job are the researchers who do it every day, with every incentive to speed up their own labs.

The early evidence is in. RE-Bench, a benchmark built to measure how well AI does AI research, found that the best AI agents scored four times higher than human experts when both had two hours per task, and generated and tested solutions more than ten times faster at a fraction of the cost. Humans still win with more time: given 32 hours, they scored twice as high. That gap is closing.

Here’s what the cycle may look like:

StageTimescaleKey driversStatusRate-limiting factor
InfrastructureYearsAI success → investment → better hardware and infrastructureOngoing; massive investmentHardware development cycle
Model development1–2 yearsHuman-led research with AI assistanceActive across major labsTraining run complexity
Data generationMonthsAI systems generating synthetic training dataBeginningData quality verification
Tool developmentDays to weeksAI systems creating their own scaffolding and toolsEarly experimentsSoftware integration time
Network self-improvementHours to weeksGroups of AI systems innovating “social” institutionsUnknownInference rate
Recursive improvementUnknownAGI or superintelligent systems improving themselves autonomouslyNot yet possibleUnknown

Source: Keep the Future Human.

Jim Fan, who leads NVIDIA’s AI agents work, puts it simply: “We are not truly done until transformers start to research the next transformer… I don’t think we are very far away from this.”

Nobody building it will slow down

“Show me the incentive and I’ll show you the outcome.” ― Charlie Munger

There is an arms race to build AGI first, and it is unlike any other in history. Even in the face of public backlash, I don’t believe it will stop on its own. Governments and companies will race ahead, with the ferocity of competition inversely proportional to the safety measures taken. Three incentives drive it.

Geopolitical supremacy. Superintelligence will be the most powerful technology, and the most powerful weapon, ever built. Those who have it stand to dominate those who don’t, a situation unseen since the nuclear era. For the US and China, the race is existential: whoever builds AGI first, then ASI, will hold a decisive strategic and military advantage.

Economic power. In 2025 alone, Apple, Amazon, Microsoft, Google and Meta committed more than $440 billion to AI. Governments pledged hundreds of billions more: over $500 billion in the US including the Stargate project, €200 billion from the EU, €109 billion from France and $100 billion from the UAE. Goldman Sachs forecast that around a trillion dollars will flow into AI, about four times the cost of Apollo and ten times the Manhattan Project. McKinsey estimates AI could add up to $23 trillion a year by 2040. A true AGI breakthrough could be the largest economic discontinuity in human history: an industrial revolution compressed into a few years. The size of the prize is real.

“If we don’t, someone else will.” Some believe AGI is inevitable, so being first is the only way to set the safeguards. Ilya Sutskever left OpenAI in 2024, unconvinced it was taking safety seriously enough, and raised at least $1 billion for Safe Superintelligence, with one goal and one product. It’s a powerful motivator, and it speeds up the very race it claims to make safer.

Together, these create a winner-takes-most race that rewards speed over caution. Stuart Armstrong, Nick Bostrom and Carl Shulman modelled this exact dynamic in 2013 and called it a race to the precipice.

The justification doesn’t hold

The only question that matters is whether it can get away from us.

OpenAI was founded because Elon Musk thought Google wouldn’t take safety seriously. Anthropic was founded because Dario Amodei and others didn’t trust OpenAI to take it seriously. Now both are racing to build superintelligence with little regard for safety, justified by the claim that China will build it first and take America’s global hegemony.

If OpenAI and Anthropic build something they believe is dangerous because “China will do it anyway,” they need to be extremely sure China will do that—build a recursively self-improving system—not just keep training models. Yet the labs themselves (looking at you, Anthropic) keep saying China is distilling or copying their models. In other words, China is keeping up by copying, not overtaking. If that’s true, a slowdown at the frontier labs slows China down too.

So why are we using “China will do it anyway” as the justification? Even if we don’t trust Beijing, why does that distrust outrank the survival of humanity? China’s posture towards AGI has arguably been more conservative than America’s. If progress continues as it is, industry, energy and settlement will begin moving off-world within the next 30 years, and the relative positions of Washington and Beijing will become a relic. Great power competition still matters. It does not outrank keeping the first self-improving systems from destroying the option value of the rest of humanity.

To those asking for the mechanism by which AI gets out of hand: the control problem has been written about for more than 60 years, starting with Norbert Wiener in 1960 and I.J. Good in 1965.

And to the frontier labs fighting for economic supremacy: AI in its current form is already going to be economically transformative. You will capture enormous rents without crossing into self-improving superintelligence. Take a breath, pause, and realise your decisions are determining the future of humanity.

The wager

Every person alive today, and everyone who will ever live, should share the same first incentive: that humanity continues with its prosperity and agency intact.

Blaise Pascal argued that you should live as if God exists. If you’re wrong, you lose a few finite pleasures. If you’re right, you gain eternity. When the downside of an action is catastrophic and irreversible, and the cost of caution is small by comparison, you don’t wait for certainty. If there’s even a 0.01% chance of self-improving AI taking over humanity, all other concerns are secondary.

We don’t have to look far back to find a time the world neglected the warning signs. During COVID-19, the people who waited for proof ended up at the mercy of whatever happened around them. Superintelligence we cannot control is that same risk magnified to infinity. We’re not wagering millions of lives. We’re wagering everyone’s, forever.

If we slow down and the danger was overstated, we lose time and revenue. If we race ahead and it wasn’t, we won’t get a second attempt.

I’ve written about what we might do about it in A few thoughts on not going extinct from AI, and at length in The Last Invention.