arXiv:2503.05812cs.CYcs.AI2025-03被引 10

为前沿AI设定不可容忍风险阈值,防范重大安全威胁。

Intolerable Risk Threshold Recommendations for Artificial Intelligence

  • 提出八类风险的可操作阈值建议,聚焦实际应用场景
  • 强调在数据不足时追求‘足够好’而非完美阈值
  • 适合政策制定者与科技企业用于前置风险管理

前沿AI模型可能在未来几年对公共安全、人权、经济稳定及社会价值构成严重威胁,风险来源包括恶意滥用、系统故障、意外连锁反应或多个模型同时失效。2024年5月韩国首尔人工智能峰会期间,16家全球AI组织签署《前沿AI安全承诺》,27个国家及欧盟宣布将定义此类风险阈值。为履行承诺,需明确并公开‘在未充分缓解情况下,模型或系统所导致的严重风险被视作不可容忍的临界点’。本文提出关键原则与考量因素,例如在能力快速演进、数据有限的情况下,应追求‘足够好’而非完美的阈值。同时针对化学、生物、辐射和核(CBRN)武器、网络攻击、模型自主性、说服与操纵、欺骗、毒性、歧视、经济社会破坏等八类风险,提出具体阈值建议,并附案例研究。目标是为政策制定者与行业领导者提供起点或补充资源,推动以预防为主的风险管理策略,实现事前防范而非事后补救。

原文摘要 · Abstract (English)

Frontier AI models -- highly capable foundation models at the cutting edge of AI development -- may pose severe risks to public safety, human rights, economic stability, and societal value in the coming years. These risks could arise from deliberate adversarial misuse, system failures, unintended cascading effects, or simultaneous failures across multiple models. In response to such risks, at the AI Seoul Summit in May 2024, 16 global AI industry organizations signed the Frontier AI Safety Commitments, and 27 nations and the EU issued a declaration on their intent to define these thresholds. To fulfill these commitments, organizations must determine and disclose ``thresholds at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable.'' To assist in setting and operationalizing intolerable risk thresholds, we outline key principles and considerations; for example, to aim for ``good, not perfect'' thresholds in the face of limited data on rapidly advancing AI capabilities and consequently evolving risks. We also propose specific threshold recommendations, including some detailed case studies, for a subset of risks across eight risk categories: (1) Chemical, Biological, Radiological, and Nuclear (CBRN) Weapons, (2) Cyber Attacks, (3) Model Autonomy, (4) Persuasion and Manipulation, (5) Deception, (6) Toxicity, (7) Discrimination, and (8) Socioeconomic Disruption. Our goal is to serve as a starting point or supplementary resource for policymakers and industry leaders, encouraging proactive risk management that prioritizes preventing intolerable risks (ex ante) rather than merely mitigating them after they occur (ex post).

AI安全风险阈值前沿模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。