arXiv:2507.16534cs.AIcs.CL2025-07被引 11

为前沿AI风险建立评估框架,划分可接受、预警与禁止三类风险区。

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

  • 用环境-威胁-能力模型分析七类前沿AI风险
  • 多数模型处于绿色或黄色风险区,未触及红色禁线
  • 适用于政策制定者和安全研究人员参考

为理解快速发展的前沿人工智能(AI)模型带来的前所未有风险,本报告基于前沿AI风险管理体系(v1.0)的E-T-C分析框架(部署环境、威胁来源、赋能能力),系统识别了七类关键风险:网络攻击、生物化学风险、说服操控、不受控的自主AI研发、战略欺骗与谋划、自我复制及合谋。依据“AI-45°定律”,通过“红灯”(不可接受阈值)和“黄灯”(预警指标)划定风险区域:绿区(可管理,常规部署并持续监控)、黄区(需加强缓解措施,受限部署)、红区(须暂停研发或部署)。实验结果表明,所有近期前沿AI模型均位于绿区或黄区,未越过红区。具体而言,无模型触碰网络攻击或不受控研发的黄灯线;自我复制与战略欺骗风险多数模型在绿区,部分推理模型位于黄区;在说服操控方面,多数模型处于黄区,因对人类影响显著;生物化学风险方面,尚无法排除多数模型处于黄区的可能性,需更深入威胁建模与评估。本工作反映了当前对前沿AI风险的理解,并呼吁集体行动应对挑战。

原文摘要 · Abstract (English)

To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier risks. Drawing on the E-T-C analysis (deployment environment, threat source, enabling capability) from the Frontier AI Risk Management Framework (v1.0) (SafeWork-F1-Framework), we identify critical risks in seven areas: cyber offense, biological and chemical risks, persuasion and manipulation, uncontrolled autonomous AI R\&D, strategic deception and scheming, self-replication, and collusion. Guided by the "AI-$45^\circ$ Law," we evaluate these risks using "red lines" (intolerable thresholds) and "yellow lines" (early warning indicators) to define risk zones: green (manageable risk for routine deployment and continuous monitoring), yellow (requiring strengthened mitigations and controlled deployment), and red (necessitating suspension of development and/or deployment). Experimental results show that all recent frontier AI models reside in green and yellow zones, without crossing red lines. Specifically, no evaluated models cross the yellow line for cyber offense or uncontrolled AI R\&D risks. For self-replication, and strategic deception and scheming, most models remain in the green zone, except for certain reasoning models in the yellow zone. In persuasion and manipulation, most models are in the yellow zone due to their effective influence on humans. For biological and chemical risks, we are unable to rule out the possibility of most models residing in the yellow zone, although detailed threat modeling and in-depth assessment are required to make further claims. This work reflects our current understanding of AI frontier risks and urges collective action to mitigate these challenges.

AI风险安全管理前沿技术风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。