arXiv:2607.16112cs.AI2026-07被引 1

统一前沿AI安全阈值,提升风险评估一致性

Harmonizing AI Safety Thresholds

论文配图:Harmonizing AI Safety Thresholds
图 1 · 摘自论文原文
  • 以预期损害为核心,建模滥用风险的传播路径与发布条件
  • 基于实际AI进展速度设定自动化研发安全阈值
  • 为监管与企业对比提供可验证的通用标准,适合政策制定者

前沿AI公司发布的性能阈值差异显著,导致第三方难以验证是否突破阈值,也难以跨公司比较安全要求。缺乏统一的最低标准可能导致风险管理不一致,形成安全标准下滑的潜在竞赛。本文提出一种跨三个风险领域(滥用风险、生物与网络攻击、自动化AI研发)的阈值协调方法。针对滥用风险,以预期损害为基本指标,采用显式风险建模,考虑风险传播路径与模型发布条件;针对自动化AI研发,基于观测到的AI进步速率而非预期危害设定阈值。分析拓展了已有研究,揭示了现有实证数据的空白与局限。

原文摘要 · Abstract (English)

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

AI安全风险评估阈值协调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。