arXiv:2412.02512cs.CYcs.AI2024-12

为提前预警高风险AI能力,提出分级信息共享框架

Pre-Deployment Information Sharing: A Zoning Taxonomy for Precursory Capabilities

  • 按危险程度划分能力阶段,建立预警区分类别体系
  • 要求在部署前即通报早期风险能力,由安全机构统一接收
  • 推动跨国安全机构共享信息,防范区域风险外溢

高影响力且潜在危险的能力应在触及红线前分解为早期预警信号。每个预警信号对应一个前期能力,其位于接近最终高影响能力的连续谱上。为有效检测与追踪能力进展,本文提出一种与分阶段信息交换框架相匹配的危险能力区分类别体系。根据前沿AI安全承诺中的第七条,签署方承诺在适当情况下向可信主体(包括指定机构)分享更详细信息。基于该分类体系,本文提出四项具体建议:(1)一旦内部评估确认存在前期能力,应立即通报;(2)人工智能安全研究所(AISIs)应作为指定接收与协调机构;(3)随着前期能力向红线推进,AISIs须建立充分的信息保护机制,必要时可对相关信息进行分类或标记为受控;(4)某一地区高影响力能力进展可能引发其他地区风险,需开展更全面的国际风险评估,因此各AISIs应依据现有国际保密交换框架,借鉴其他高风险行业经验,相互交换前期能力信息。

原文摘要 · Abstract (English)

High-impact and potentially dangerous capabilities can and should be broken down into early warning shots long before reaching red lines. Each of these early warning shots should correspond to a precursory capability. Each precursory capability sits on a spectrum indicating its proximity to a final high-impact capability, corresponding to a red line. To meaningfully detect and track capability progress, we propose a taxonomy of dangerous capability zones (a zoning taxonomy) tied to a staggered information exchange framework that enables relevant bodies to take action accordingly. In the Frontier AI Safety Commitments, signatories commit to sharing more detailed information with trusted actors, including an appointed body, as appropriate (Commitment VII). Building on our zoning taxonomy, this paper makes four recommendations for specifying information sharing as detailed in Commitment VII. (1) Precursory capabilities should be shared as soon as they become known through internal evaluations before deployment. (2) AI Safety Institutes (AISIs) should be the trusted actors appointed to receive and coordinate information on precursory components. (3) AISIs should establish adequate information protection infrastructure and guarantee increased information security as precursory capabilities move through the zones and towards red lines, including, if necessary, by classifying the information on precursory capabilities or marking it as controlled. (4) High-impact capability progress in one geographical region may translate to risk in other regions and necessitates more comprehensive risk assessment internationally. As such, AISIs should exchange information on precursory capabilities with other AISIs, relying on the existing frameworks on international classified exchanges and applying lessons learned from other regulated high-risk sectors.

AI安全风险预警信息共享

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。