重审前沿AI安全论证基础,构建更可靠的安全保障框架。
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
- 借鉴航空核能等领域的安全论证方法,重构AI对齐安全案例
- 指出当前对齐社区方法存在理论缺陷,难以支撑可信安全论证
- 以欺骗性对齐和生化武器能力为例,提出可落地的安全论证范式
本文探讨前沿AI系统安全论证的理论基础。安全论证是结构化、可辩护的论据,用于证明系统在特定场景下可接受地安全。该方法已广泛应用于航空航天、核能和汽车等高危领域。随着生成式AI领导者提出的《新加坡全球AI安全研究优先事项共识》和《国际AI安全报告》,前沿AI的安全论证日益受到重视。本文评估现有工作,指出对齐社区虽借鉴了保障领域经验,但存在显著局限。为此,本文重新思考对齐安全论证方法,从安全保证领域汲取成熟理论与方法,分析其在对齐社区应用中的不足。基于此,以欺骗性对齐和生化放射核(CBRN)能力为案例,结合现有理论性安全论证草图,提出一个可操作的安全论证框架。本研究通过严谨的理论与方法,为构建稳健、可辩护且实用的安全论证体系提供基础支持,助力保障前沿AI系统的安全性。
原文摘要 · Abstract (English)
This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, they have been used in safety-critical industries, such as aerospace, nuclear or automotive. As a result, safety cases for frontier AI have risen in prominence, both in the safety policies of leading frontier developers and in international research agendas proposed by leaders in generative AI, such as the Singapore Consensus on Global AI Safety Research Priorities and the International AI Safety Report. This paper appraises this work. We note that research conducted within the alignment community which draws explicitly on lessons from the assurance community has significant limitations. We therefore aim to rethink existing approaches to alignment safety cases. We offer lessons from existing methodologies within safety assurance and outline the limitations involved in the alignment community's current approach. Building on this foundation, we present a case study for a safety case focused on Deceptive Alignment and CBRN capabilities, drawing on existing, theoretical safety case "sketches" created by the alignment safety case community. Overall, we contribute holistic insights from the field of safety assurance via rigorous theory and methodologies that have been applied in safety-critical contexts. We do so in order to create a better foundational framework for robust, defensible and useful safety case methodologies which can help to assure the safety of frontier AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。