arXiv:2509.24394cs.CYcs.AI2025-09

用行为分析法揭示OpenAI安全框架实则允许高危部署

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies

  • 基于可用性理论解析框架条款,识别实际允许与禁止的行为
  • 仅覆盖少数AI风险,却允许部署可能致超千人死亡的系统
  • 适合关注AI治理漏洞或政策评估的研究者阅读

头部人工智能公司推出'安全框架'作为自愿性自我监管工具,宣称设定风险阈值与安全流程以规范高度智能AI的研发与部署。理解这些声明涵盖哪些风险、允许或禁止何种行为,对评估其实际治理效果至关重要。本文运用可用性理论,结合机制与条件模型及MIT AI风险库,分析OpenAI《2025年准备就绪框架》(版本2)。结果发现:该框架仅要求评估极小部分AI风险;鼓励部署具备'中等'能力但可能意外引发'严重危害'(定义为死亡人数超1000人或损失超1000亿美元)的系统;且允许其首席执行官部署更危险的能力。研究表明,当前行业自律难以有效缓解AI风险,亟需更有力的治理干预。本研究提供的可用性分析方法可复现,用于评估安全框架的真实允准范围。

原文摘要 · Abstract (English)

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly capable AI. Understanding which AI risks are covered and what actions are allowed, refused, demanded, encouraged, or discouraged by these statements is vital for assessing how these frameworks actually govern AI development and deployment. We draw on affordance theory to analyse the OpenAI 'Preparedness Framework Version 2' (April 2025) using the Mechanisms & Conditions model of affordances and the MIT AI Risk Repository. We find that this safety policy requests evaluation of a small minority of AI risks, encourages deployment of systems with 'Medium' capabilities for unintentionally enabling 'severe harm' (which OpenAI defines as >1000 deaths or >$100B in damages), and allows OpenAI's CEO to deploy even more dangerous capabilities. These findings suggest that effective mitigation of AI risks requires more robust governance interventions beyond current industry self-regulation. Our affordance analysis provides a replicable method for evaluating what safety frameworks actually permit versus what they claim.

AI治理安全框架可用性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。