arXiv:2503.18956cs.CYcs.AI2025-03综述被引 7

提出可落地的全球AI安全条约,用算力阈值管控高风险模型开发。

International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty

  • 设定算力门槛,超限模型需接受国际安全审计。
  • 建立跨国AI安全研究所网络,有权叫停高风险研发。
  • 兼顾可操作性与灵活性,适合政策制定者参考。

先进通用人工智能(GPAI)的恶意使用或故障可能引发人类文明被边缘化甚至灭绝的风险,据顶尖专家评估。为应对这一挑战,近年涌现出大量关于国际AI安全协议的提案。本文回顾2023年以来的相关提议,梳理共识与分歧,并结合既有文献评估其可行性。重点探讨风险阈值、监管方式、协议类型及五大关键进程:科学共识构建、标准化、审计、验证与激励机制。基于此,我们提出一项条件性AI安全条约:当模型训练算力超过预设阈值时,必须接受严格监督;同时要求对模型、信息安全和治理实践进行互补性审计,由国际AI安全研究所(AISIs)网络负责,具备在风险不可接受时暂停开发的权力。该方案融合即刻可实施措施与可适应持续研究的灵活结构。

原文摘要 · Abstract (English)

The malicious use or malfunction of advanced general-purpose AI (GPAI) poses risks that, according to leading experts, could lead to the 'marginalisation or extinction of humanity.' To address these risks, there are an increasing number of proposals for international agreements on AI safety. In this paper, we review recent (2023-) proposals, identifying areas of consensus and disagreement, and drawing on related literature to assess their feasibility. We focus our discussion on risk thresholds, regulations, types of international agreement and five related processes: building scientific consensus, standardisation, auditing, verification and incentivisation. Based on this review, we propose a treaty establishing a compute threshold above which development requires rigorous oversight. This treaty would mandate complementary audits of models, information security and governance practices, overseen by an international network of AI Safety Institutes (AISIs) with authority to pause development if risks are unacceptable. Our approach combines immediately implementable measures with a flexible structure that can adapt to ongoing research.

AI安全国际条约算力监管治理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。