arXiv:2505.17084cs.CRcs.AI2025-05

借鉴核电安全经验,用非概率方法提升大模型系统安全性

From nuclear safety to LLM security: Applying non-probabilistic risk management strategies to build safe and secure LLM-powered systems

  • 引入核工程等领域的非概率风险策略应对大模型安全挑战
  • 提出5类100多种可复用的安全设计方法,适配对抗性攻击场景
  • 适合系统架构师和负责任AI实践者参考落地

大语言模型(LLMs)虽能力强大,但其安全与安全挑战难以用传统概率风险分析(PRA)应对。由于系统新颖性和复杂性,尤其是面对自适应对手时,全面枚举和量化风险不切实际。借鉴核能、土木工程等领域通用的非概率风险管理策略(如事件树分析、鲁棒设计),本文提出超过100种适用于大模型系统的非概率风险应对方法,并将其划分为五大类别,映射至大模型安全与人工智能安全范畴。同时构建了基于大模型的自动化应用流程,支持解决方案架构师采用这些策略。尽管存在局限,这些方法仍可为大模型系统的安全性、可靠性及负责任人工智能提供有效支撑。

原文摘要 · Abstract (English)

Large language models (LLMs) offer unprecedented and growing capabilities, but also introduce complex safety and security challenges that resist conventional risk management. While conventional probabilistic risk analysis (PRA) requires exhaustive risk enumeration and quantification, the novelty and complexity of these systems make PRA impractical, particularly against adaptive adversaries. Previous research found that risk management in various fields of engineering such as nuclear or civil engineering is often solved by generic (i.e. field-agnostic) strategies such as event tree analysis or robust designs. Here we show how emerging risks in LLM-powered systems could be met with 100+ of these non-probabilistic strategies to risk management, including risks from adaptive adversaries. The strategies are divided into five categories and are mapped to LLM security (and AI safety more broadly). We also present an LLM-powered workflow for applying these strategies and other workflows suitable for solution architects. Overall, these strategies could contribute (despite some limitations) to security, safety and other dimensions of responsible AI.

大模型安全风险评估非概率方法AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。