arXiv:2508.13700cs.CYcs.AI2025-08被引 7

系统梳理从个体伤害到人类存亡的AI风险谱系,揭示潜在灾难路径。

The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats

  • 按滥用、错位、系统性三类归因,构建风险框架
  • 指出当前行为已现权力集中与人类能力退化苗头
  • 强调竞争与协作缺失会加剧所有风险,需主动应对

随着人工智能系统日益强大、融合与普及,理解其伴随的风险愈发重要。本文绘制了从影响个体用户当前伤害到威胁人类生存的极端风险全谱。将风险分为三大因果类别:滥用风险(如用AI制造生物武器、发动网络攻击或部署致命自主武器)、错位风险(系统追求与人类价值观冲突的目标,包括奖励作弊、策略谋划与权力欲望)以及系统性风险(如权力集中、政治经济去权、人类依赖过度导致能力退化,或固化现有价值阻碍未来道德进步)。此外,识别出竞争压力、事故、企业冷漠与协调失败等风险放大因素,使各类风险更易发生且更严重。文章将当前可观察的AI行为与未来可能的灾难性后果相联系,表明现有趋势若不加干预,可能演化为灾难性结局。我们相信美好未来仍可实现,但不会自动发生;唯有前所未有的协同努力,才能通向非凡未来。

原文摘要 · Abstract (English)

As AI systems become more capable, integrated, and widespread, understanding the associated risks becomes increasingly important. This paper maps the full spectrum of AI risks, from current harms affecting individual users to existential threats that could endanger humanity's survival. We organize these risks into three main causal categories. Misuse risks, which occur when people deliberately use AI for harmful purposes - creating bioweapons, launching cyberattacks, adversarial AI attacks or deploying lethal autonomous weapons. Misalignment risks happen when AI systems pursue outcomes that conflict with human values, irrespective of developer intentions. This includes risks arising through specification gaming (reward hacking), scheming and power-seeking tendencies in pursuit of long-term strategic goals. Systemic risks, which arise when AI integrates into complex social systems in ways that gradually undermine human agency - concentrating power, accelerating political and economic disempowerment, creating overdependence that leads to human enfeeblement, or irreversibly locking in current values curtailing future moral progress. Beyond these core categories, we identify risk amplifiers - competitive pressures, accidents, corporate indifference, and coordination failures - that make all risks more likely and severe. Throughout, we connect today's existing risks and empirically observable AI behaviors to plausible future outcomes, demonstrating how existing trends could escalate to catastrophic outcomes. Our goal is to help readers understand the complete landscape of AI risks. Good futures are possible, but they don't happen by default. Navigating these challenges will require unprecedented coordination, but an extraordinary future awaits if we do.

AI风险伦理治理系统性风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。