总结前沿AI安全框架的最新实践,助力系统性防范重大风险
Emerging Practices in Frontier AI Safety Frameworks
- 聚焦风险识别、缓解与治理三大核心领域
- 提炼多方协同制定的安全框架新兴做法
- 为开发者与政策制定者提供可参考的实践指南
作为2024年韩国首尔峰会达成的前沿AI安全承诺的一部分,众多AI开发者同意发布安全框架,阐明其如何管理其系统可能带来的严重风险。本文总结了企业、政府及研究人员当前关于如何制定有效安全框架的共识。我们梳理了安全框架的三个核心组成部分:风险识别与评估、风险缓解措施、治理机制,并在每个领域中识别出新兴实践。由于安全框架尚属新兴且快速发展,本文旨在成为迄今工作的概览,并为后续讨论与创新提供起点。
原文摘要 · Abstract (English)
As part of the Frontier AI Safety Commitments agreed to at the 2024 AI Seoul Summit, many AI developers agreed to publish a safety framework outlining how they will manage potential severe risks associated with their systems. This paper summarises current thinking from companies, governments, and researchers on how to write an effective safety framework. We outline three core areas of a safety framework - risk identification and assessment, risk mitigation, and governance - and identify emerging practices within each area. As safety frameworks are novel and rapidly developing, we hope that this paper can serve both as an overview of work to date and as a starting point for further discussion and innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。