A few thoughts on not going extinct from AI
Everyone benefits from safer AI. Whoever funds the research captures only a fraction of that benefit.
Whoever solves AI alignment—assuming it can be solved—will win the Nobel Peace Prize of Nobel Peace Prizes.
Last month, Jacob Coxon dropped a proverbial bomb on Twitter, warning that the frontier AI labs had little regard for safety and were stuck in a race dynamic that could lead to human extinction if left undisturbed and unregulated.
This week, the American race participants sat down at the White House and signed an accord, which the president labelled “morally binding,” to police themselves.
This is a real step in the right direction, but it doesn’t make the problem of AI alignment go away. The accord is about how the labs conduct themselves—controls, audits, board oversight—whereas alignment is still an open research question. That is, we don’t yet know how to verify that a system smarter than us is safe, so there’s a limit to what an auditor can actually check.
Further, I’d wager a voluntary agreement only holds for as long as it’s cheap to keep. Each lab still has the same incentive to move faster that it had last week, and nothing in the accord rewards anyone for actually solving the problem.
People differ on the odds of misaligned AI leading to human extinction, but given the stakes of getting this wrong, even if there’s a 0.01% chance of us losing control of AI, we need to take alignment extremely seriously.
Further, I’m hopeful that China and the US can co-operate here. If our technological progress continues the way it’s going, in the next 30 years, we will likely have industry, energy, and settlement move off-world, and Washington versus Beijing in relative positions will increasingly be a relic of the past, as we explore the final frontier of our galaxy.
Great power competition still matters, but it does not outrank keeping the first self-improving systems from destroying the option value for the rest of humanity.
The question, to me, is: given we do not yet have a solution to aligning superintelligent systems, how might we mobilise the best of human ingenuity to solve this?
Alignment has the economics of a global public good: everyone benefits from making AI safer, but whoever pays for the research (right now, primarily the AI labs) captures only a fraction of that benefit. Secondly, the labs are now slowed because of their need to solve alignment issues before releasing more advanced models that would invariably allow them to keep leading the race.
So two things now seem true to me:
- For the first time, the labs have an existential (company survival) reason to solve alignment, and will begin wearing the full cost of this.
- Their incentive to have others help solve alignment has just increased by orders of magnitude, because solving alignment allows them to release advanced models and retain market share.
So how can we help the labs help humanity? What can we do to catalyse the wisdom of the crowd and allow more teams and individuals to contribute to alignment? I think it starts with two basic principles:
1) Make the incentives strong enough
Here, I’m primarily thinking of what an X-prize for solving alignment would look like. That is, if you solve alignment, you receive a pot of gold.
I’d wager the amount of capital willing to donate to this would be at least hundreds of billions, taking into account philanthropic, public, and private capital (including the labs themselves and lab employees).
Given we have little knowledge of how we’re going to solve alignment, the problem should be treated as basic research. That is, we need to fund many competing approaches at the outset with the knowledge that most will fail.
China is an interesting case study in funding competing approaches. In strategic industries, it supports firms through cheap credit, tax breaks, subsidised land, and direct grants. The principle worth borrowing is funding multiple approaches before we know which will work.
I’m skeptical traditional venture capital can solve this problem given the tragedy of the commons problem mentioned above, i.e., the returns to solving alignment are too diffuse to be captured by private market investors.
We could create milestones where teams are paid for intermediate progress, where we incentivise as much knowledge sharing as possible between competing teams, and so on—this would ensure we have something closer to a market price.
The specific model should be debated, but if capitalism got us here, it can likely get us out.
2) Open the frontier enough for anyone working on alignment to have access to leading AI models
Alongside creating adequate incentives, we would have to ensure that participants have access to each other’s research—the most fundamental of science principles. There are encouraging moves in this direction, with examples like Base Labs, who are at the frontier of AI research and committed to open source.
But overall, we’re currently preventing the wisdom of the crowd from contributing by having a fully closed frontier.
However, we need not require the labs to be completely open. We might just require a collective of teams participating to adhere to strict, verifiable information sharing regimes—similar to how an investment bank creates ‘Chinese walls’ between its trading desk and M&A division. In principle, any unauthorised sharing would be punishable by law, and sharing would be monitored on device.
Once again, the specific model should be debated, but if the basic principles of science got us here, a return to them might just get us out.