为复杂AI系统设计可量化的风险评估框架,提升安全预判能力。
Adapting Probabilistic Risk Assessment for AI
- 基于高可靠性行业经验,构建AI风险的系统化分析方法。
- 通过因果链建模识别从技术特性到社会影响的风险路径。
- 支持开发者、评估者和监管方进行透明、可追溯的风险决策。
现代通用人工智能系统因能力快速演化及潜在重大危害,带来紧迫的风险管理挑战,现有方法常依赖选择性测试与未明示的假设,未能全面评估AI对社会与生物圈造成的直接或间接风险路径。本文提出面向AI的概率风险评估(PRA)框架,借鉴核电、航天等高可靠性行业的成熟技术,指导评估者识别风险、估算可能性与严重程度区间,并显式记录证据、假设与分析粒度。该框架的工具将结果整合为风险报告卡,包含所有评估风险的聚合估计。其三大创新包括:(1)面向特性的危害分析,基于人工智能系统要素(如能力、领域知识、可操作性)的一阶分类体系实现系统覆盖;(2)风险路径建模,采用双向分析与前瞻性技术刻画从系统特性到社会影响的因果链;(3)不确定性管理,通过情景分解、参考尺度与明确追踪协议,构建在新颖或数据有限条件下的可信预测。框架还通过整合证据,将不同评估方法统一为可比的绝对风险量化值,支持全生命周期决策。该框架已实现为供开发者、评估者与监管方使用的交互式工作簿工具。
原文摘要 · Abstract (English)
Modern general-purpose artificial intelligence (AI) systems present an urgent risk management challenge, as their rapidly evolving capabilities and potential for catastrophic harm outpace our ability to reliably assess their risks. Current methods often rely on selective testing and undocumented assumptions about risk priorities, frequently failing to make a serious attempt at assessing the set of pathways through which AI systems pose direct or indirect risks to society and the biosphere. This paper introduces the probabilistic risk assessment (PRA) for AI framework, adapting established PRA techniques from high-reliability industries (e.g., nuclear power, aerospace) for the new challenges of advanced AI. The framework guides assessors in identifying potential risks, estimating likelihood and severity bands, and explicitly documenting evidence, underlying assumptions, and analyses at appropriate granularities. The framework's implementation tool synthesizes the results into a risk report card with aggregated risk estimates from all assessed risks. It introduces three methodological advances: (1) Aspect-oriented hazard analysis provides systematic hazard coverage guided by a first-principles taxonomy of AI system aspects (e.g. capabilities, domain knowledge, affordances); (2) Risk pathway modeling analyzes causal chains from system aspects to societal impacts using bidirectional analysis and incorporating prospective techniques; and (3) Uncertainty management employs scenario decomposition, reference scales, and explicit tracing protocols to structure credible projections with novelty or limited data. Additionally, the framework harmonizes diverse assessment methods by integrating evidence into comparable, quantified absolute risk estimates for lifecycle decisions. We have implemented this as a workbook tool for AI developers, evaluators, and regulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。