There is a high chance of achieving (1) given the competitive dynamics between geopolitical rivals in particular.
(2) might occur in the absence of kill switches or neglect in cybersecurity relating to GPU weights.
(3) is too tenuous at this stage. Both ends of the literature suggest no one really knows what AI values are, nor what mechanisms and enforcement designs contain them. Yes we have misalignment, sandboxing errors, and reward hacking. Yet this does not imply malicious objectives from AI as a "second species" will arise in the future, although this cannot be ruled out: particularly with RSI causing an intelligence explosion as Turing predicted. However, if it becomes difficult for us humans to control them, then automating AI research could become especially socially valuable, which was why the labs racing was fully consistent with them taking AI-safety and X-risks seriously (their self-harming, in valuations, calls to pace the frontier are the best evidence possible for this). It remains the case, as the most likely X-risk channels highlighted here suggest, that use by malicious human actors pose the greater risks.
(4) however was demonstrated with Hugging Face. This therefore raises my p(doom) slightly.
On the other hand, Yudkowsky and the doomers suggest that constitutions have completely failed in alignment. The theory and data do NOT support that. The evolutionary game dynamics, and the AI obsession with pure math (which by definition has no real-world application, and is primarily pursued for beauty, and manifolds or group theory are certainly beautiful!), suggest a harmful second species colluding against us to wipeout humanity is not an inevitability, nor even likely.
This makes me fairly optimistic that, with sufficient talent directed towards advancing AI-safety, that we'll solve the alignment problem once RSI occurs. However, this is not guaranteed either. The labs are making the biggest bet that humanity has taken by far; that RSI will automate and solve the alignment problem, and the ensuing intelligence explosion will solve most of humanity's suffering.
My everchanging p(doom) probably sits somewhere just over halfway (to account for the right-tailed distribution here) between my pre and (initial) post Hugging-Face likelihoods: perhaps 8%. It would be unwise not to update since this saga unfolded.
Some of the latest macroeconomic models from the likes of Anthropic and Acemoglu suggest rather grim futures for employment relative to what economists were previously predicting, hence its fair to say our methods lean overly conservative towards excess rigour here. Perhaps the recent backlash from the mathematicians, and some of the sentiment behind the anti data-centre movement, reflect these (increasingly justified) anxieties?
Nonetheless, this is orders of magnitude more dangerous than nuclear weapons. That's enough for me to take AI-safety incredibly seriously!

