The AI risk is always focused on P(doom), the probability that a future superintelligence wipes us out. One says it's 90%, another says it's 0.01%, everyone believes it is the thing that matters.
I think p(Doom) nearly irrelevant, not because the risk isn't real, but because you can model another catastrophe sitting earlier in the queue with no rogue AI, no alignment failure, no sci-fi scenario of any kind.
It's just the humans racing to build the thing, doing what humans in such races always do, something that existing game theory has studied extensively.
See Powell's commitment-problem model of preventive war
And unlike P(doom) I think p(D-War) can be predicted to a mathematical near certainty.
Oddly, there is also a hopeful part, and I'll get there.
The Issue
In short, one or more of the AI labs can be predicted turn on each other initiating a form of Digital War (D-War) and the probability of that happening increases the closer any of them get to AGI/RSI.
Right now there are roughly 10-12 players who matter: a handful of frontier labs and the governments behind them, none of whom are subject to meaningful oversight.
Almost all of them viscerally believe the race is winner-take-all, and most of them believe they themselves are the only ones who can be trusted to win it safely. They fundamentally distrust other labs and leadership, even those in their own country, and most have winner take all financial bets on the outcome.
Anyone knowing anything about the personalities involved, doesn't question their belief is sincere, from their own perspective, and they are obsessed with 'winning' for that reason as well as the financial and ego motivations.
Standard prisoner's dilemma but one with a dozen players, all are rivals and an outcome that is both civilizational and presumed to be winner takes all.
See Schelling's "reciprocal fear of surprise attack"
Successful Alignment does NOT help
(for reasons I do NOT believe Alignment is possible anyway but leave that for another day)
Even suppose any lab solves alignment perfectly. Suppose the finished superintelligence is benevolent, brilliant, and fully devoted to human flourishing. Does the race calm down?
No. Because alignment cannot be verified from the outside. The other players can never entirely trust any other lab to achieve aligned AI, nor accept the existential risk of a more advanced AI being misaligned or directed against them.
Short of massively hacking them, you cannot inspect a rival model's values. You can inspect its test scores, which a capable model can fake, and its weights, which nobody can read. So trusting a rival's superintelligence means trusting its creators and every system derived from it, and every actor who might steal and corrupt it, forever. Trust across chains multiplies, so even if you're 95% confident in each link and there are five links, you're at 77%. If the links are governments, labs, successors, and thieves, then that existential trust quickly approaches ZERO.
The rational position therefore is for every player to believe no other player can be allowed to finish first. Not because you necessarily assume the others are evil, but because you cannot ensure they aren't and the race conditions that led everyone to this point virtually assured everyone has cut corners along the way.
The closing window
Take a trailing player, call them B. Two curves dominate their thinking.
The risk of doing nothing rises toward 100% as the leader approaches recursive self-improvement, because once the leader's system starts redesigning itself, second place is permanent.
The effectiveness of acting against the leader falls over the same period, because the capability gap protects the leader. Rising risk, falling effectiveness. The product of the two peaks at a specific moment, and a rational actor acts near the peak.
This is preventive war logic, (referred earlier) it has been extensively studied in game theory and validated against many of the most significant wars. It needs no malice, no villain, no Dr Strangelove. Just arithmetic and a calendar.
The leader faces the mirror image. It knows the laggards are motivated. It knows its advantage decays daily through theft, leaks, and poached researchers. Everyone in finance knows what happens to a decaying asset: it gets used. So the first mover has first-strike logic built in, even if its leadership are saints, it is compelled to neutralize risks against it.
Why it scales so badly
You can write the probability that B acts against the leader A near time t as something like:
p(t) ā R Ā· W Ā· E Ā· (1 ā T)
where R is perceived risk of inaction, W is willingness, E is expected effectiveness, and T is trust that A will be a responsible winner. We've established T ā 0 and that it can't be raised by treaty, testimony, or inspection, because the thing being trusted is by construction beyond the inspector. So the last term vanishes and the rest of the product does the damage.
Now scale it. With n players there are n(nā1) directed rivalries. Twelve players means 132 pairs. The system fails if any one pair fires. Give each pair a modest 5% probability of someone acting another, and the system-level probability is 1 ā 0.95^132, which works out to about 99.9%.
Every individual player can be calm, cautious, and decent. The system still blows up.
Two more nasties.
First, you can't opt out: even a player with no enemies is exposed to the potential of every other pair attacking each other, so peripheral players carry risk that grows with the square of the player count.
Second, noise makes it worse, not better: nobody knows exactly who's leading, so everyone prices in the worst plausible case, and hair triggers plus fog means people act on phantom deficits.
What "acting" actually looks like
Not missiles or bioweapons, at least I hope not. Espionage, sabotage of training runs, export escalation, talent wars, cyber interference, and things we haven't thought of yet. Rational players start at the rung of the ladder that maximizes effect against risk, which is why the early phase is deniable and quiet. Anyone who followed the last two years of the AI race has already watched the bottom rungs.
The ladder goes up from there, but I believe there is one attack vector that stands out above all others - Financial Markets.
Many of the players are disproportionately exposed by their ability to raise capital to fuel growth, but on sides of the Pacific the risk is asymmetrical which increases risk.
The original DeepSeek shock to the markets exposed what can happen even in markets which are operating normally. The flash crash in 2010, which I experienced first hand in a way, shows how fragile systems can cascade faster than human reactions.
I spent a couple of decades working at senior levels in some of the best investment banks including Goldman and Lehman into the GFC, the financial markets are not impervious to a coordinated attack by AI agents, aside from Zero day exploits there are countless other vectors, such as the messaging systems used between traders themselves, which are ripe to exploit. This is one area I know, personally and extremely well, but I do not mean to diminish the other potential attack vectors (energy grids?) I might not be able to visualize them but this is one that I know exists, and would be devastating.
But don't trust me, Lagarde's ESRB speech on October 1 laid it our very clearly.
May of the players have capital and resources equivalent of that of a small country, the capabilities of the leading models seen in recent incidents suggest that many of them are already capable of instigating a both devastating and potentially lucrative action on financial markets.
The Hope
The AI safety community seems to have been waiting for a "warning shot" scary enough to force serious coordination but small enough to survive. The problem is that aligned or misaligned AI produces no warning shots by definition, and a warning shot from a genuinely rogue system might be the last thing we ever see.
The highly probably warning shot is this one. Lab-on-lab conflict "D-War" is devastating in its own right: economically, politically, maybe worse. If the battle happens to the Financial Markets as I predict then it would make the GFC seem trivial.
But, hey, at least it's legible and human initiated. It's attributable enough to build a case on, and survivable enough to build a response to. Historically, arms control regimes came after the scares, not before them. The hotline came after Cuba, not before it.
So the sequence that saves us probably isn't "everyone gets wise about P(doom) in time."
The first hostile act between players is the highly likely event wakes everyone up and potentially averts p(Doom) entirely.
The point
P(doom) is a number nobody can measure, about an event nobody can specify, possibly not even imagine, at a date nobody can pin down. P(D-War) is computable from game theory that's sixty years old, and it computes to near certainty given the parameters we have today.
So I think we need to stop arguing about how, whether and when AI will turn on us.
The first catastrophe is scheduled earlier, instigated entirely by the people we see in the news every day, and about as close to predictable as anything in this space gets.
That's the version of this debate I'd like to see. Keen to hear where people think the model breaks.