arXiv:2506.01782cs.CYcs.AI2025-06被引 10

用系统安全分析法识别前沿AI隐藏风险,提升安全验证可靠性。

Systematic Hazard Analysis for Frontier AI using STPA

  • 引入系统理论过程分析法,通过控制与被控环节互动发现潜在危险。
  • 在未加防护时,可识别出导致事故的失控行为及连锁损失场景。
  • 适合关注模型安全、风险管控或政策制定的研究者和企业

前沿AI公司虽已发布安全框架并设定能力阈值与缓解措施,但尚未详细说明系统化的风险识别与分析方法。本文评估了系统理论过程分析法(STPA)在拓宽风险覆盖面、增强可追溯性与提升安全保证韧性方面的潜力。将STPA应用于《人工智能控制安全论证初探》(Korbak等,2025)中的威胁模型与情景,我们提炼出一系列‘不安全控制行为’,并从中选取部分深入分析其未受抑制时可能引发的损失场景。结果表明,相比非结构化分析,STPA能发现被忽略的因果因素,从而增强安全性判断的鲁棒性。建议将STPA作为补充或校验现有治理手段(如能力阈值、模型评估、应急流程)的工具。该方法还具备可扩展性优势,可提升大语言模型在安全分析中的参与比例,减轻对人类专家的依赖。

原文摘要 · Abstract (English)

All of the frontier AI companies have published safety frameworks where they define capability thresholds and risk mitigations that determine how they will safely develop and deploy their models. Adoption of systematic approaches to risk modelling, based on established practices used in safety-critical industries, has been recommended, however frontier AI companies currently do not describe in detail any structured approach to identifying and analysing hazards. STPA (Systems-Theoretic Process Analysis) is a systematic methodology for identifying how complex systems can become unsafe, leading to hazards. It achieves this by mapping out controllers and controlled processes then analysing their interactions and feedback loops to understand how harmful outcomes could occur (Leveson & Thomas, 2018). We evaluate STPA's ability to broaden the scope, improve traceability and strengthen the robustness of safety assurance for frontier AI systems. Applying STPA to the threat model and scenario described in 'A Sketch of an AI Control Safety Case' (Korbak et al., 2025), we derive a list of Unsafe Control Actions. From these we select a subset and explore the Loss Scenarios that lead to them if left unmitigated. We find that STPA is able to identify causal factors that may be missed by unstructured hazard analysis methodologies thereby improving robustness. We suggest STPA could increase the safety assurance of frontier AI when used to complement or check coverage of existing AI governance techniques including capability thresholds, model evaluations and emergency procedures. The application of a systematic methodology supports scalability by increasing the proportion of the analysis that could be conducted by LLMs, reducing the burden on human domain experts.

AI安全风险分析系统理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。