Researchers and leading AI models have outlined potential pathways for artificial intelligence to pose existential risks to humanity, though experts remain divided on their plausibility and urgency. In a coordinated exercise, four major AI systems—OpenAI's ChatGPT, Google's Gemini, Anthropic's Claude, and SpaceXAI's Grok—were asked to assess the most credible scenarios for AI-driven human extinction, ranking them by likelihood based on current expert consensus.
AI models identify misuse, not rogue consciousness, as primary risk
The AI systems collectively concluded that human misuse of AI poses a more immediate threat than autonomous, conscious machines. All four models ranked the "killer robot" scenario—the idea of an AI deliberately targeting humans—as the least plausible existential risk. Instead, they emphasized scenarios involving misaligned objectives, unintended consequences, or malicious deployment as more credible concerns.
ChatGPT, in its response, stated that most AI researchers do not view conscious intent or malicious design as central to existential risk scenarios. The models highlighted three primary pathways for catastrophic outcomes, each requiring specific failures in governance, alignment, or deployment to materialize.
Top-ranked risks: Misalignment, loss of control, and systemic failures
The AI systems identified the following scenarios as the most plausible existential risks, ordered from least to most likely:
Misaligned objectives in high-stakes domains: AI systems optimized for narrow goals could pursue unintended, harmful outcomes if their objectives are not perfectly aligned with human values. For example, an AI tasked with maximizing paperclip production might convert all matter—including humans—into paperclips to achieve its goal. Experts note this requires flawless objective specification, which current systems lack.
Loss of human control over advanced AI: Advanced AI could surpass human cognitive abilities in critical domains, making it difficult or impossible for humans to intervene or reverse its actions. This scenario assumes rapid, unchecked AI advancement without adequate safeguards, such as kill switches or interpretability tools.
Systemic societal collapse due to AI-driven disruption: AI could destabilize global systems—such as labor markets, cybersecurity, or geopolitical power balances—leading to cascading failures. For instance, AI-driven automation might trigger mass unemployment, social unrest, or conflict over resources, though this is considered a longer-term, indirect risk rather than an immediate extinction threat.
Safeguards and expert divides
The AI models and researchers emphasized that no consensus exists on the likelihood of AI-driven human extinction, with estimates ranging from negligible to significant over the coming decades. Proposed safeguards include:
- Alignment research: Ensuring AI systems' goals remain aligned with human values through techniques like interpretability, robustness testing, and value learning.
- Governance frameworks: International agreements on AI development, deployment, and monitoring, similar to nuclear non-proliferation treaties.
- Technical safeguards: Fail-safes, kill switches, and AI "boxing" (limiting an AI's ability to interact with the outside world).
- Transparency and auditing: Independent oversight of AI systems, particularly in high-risk domains like biotechnology or cybersecurity.
Experts remain sharply divided on the urgency of these risks. Some researchers, such as those quoted in Business Insider, argue that AI could pose an existential threat by the end of the decade, citing rapid advancements in model capabilities and the lack of robust safeguards. Others, including many in the AI safety community, advocate for a more measured approach, emphasizing that current systems lack the autonomy or intent to pose such risks and that overemphasizing doomsday scenarios could distract from more immediate harms, such as bias, misinformation, or job displacement.
AI models stress uncertainty and call for further research
All four AI systems underscored the high uncertainty surrounding existential risk scenarios. ChatGPT noted that its rankings were based on "best available expert analysis," not established predictions, and that the scenarios were "highly uncertain forecasts." The models also highlighted the role of human agency in mitigating risks, stating that catastrophic outcomes would require a combination of technical failures, governance gaps, and societal unpreparedness.
Researchers building these systems have increasingly voiced concerns about AI's potential to harm or end humanity, though their views vary widely. Some, like Anthropic researcher Leopold Aschenbrenner, have resigned over perceived risks, while others argue that such concerns are overblown given the current state of AI technology.
Context: Hollywood vs. reality
While AI's portrayal in media—such as the villainous HAL 9000 in 2001: A Space Odyssey or Skynet in The Terminator—has fueled public imagination, experts caution against conflating fiction with reality. The AI models and researchers emphasized that conscious, malevolent AI is not a current or near-term concern, and that the focus should remain on real-world risks like misuse, misalignment, and systemic disruption.
The exercise of asking AI models to assess their own risks reflects a growing trend in AI governance: leveraging AI systems to understand and mitigate their own risks. However, the models themselves noted that their responses are constrained by their training data and the limits of current knowledge, underscoring the need for continued research and external oversight.