arXiv:2412.08653cs.CYcs.AI2024-12被引 7

AI评估能测下限风险,但无法判断上限能力与自主系统威胁。

What AI evaluations for preventing catastrophic risks can and cannot do

  • 通过充分努力可评估特定滥用风险和能力下限
  • 无法确定模型能力上限或可靠预测未来性能
  • 适合关注安全边界与治理框架的研究者

AI评估是当前防止灾难性风险安全论证的重要组成部分。本文探讨了这类评估的边界:它们能够建立模型能力的下限,并在评估者投入足够努力时评估特定滥用风险。然而,评估在现有范式下存在根本局限——无法确立能力上限、可靠预测未来模型能力,也无法稳健评估自主AI系统的风险。这意味着评估虽有价值,但不应作为确保AI安全的主要手段。文章最后提出对前沿AI安全的渐进改进建议,同时承认这些根本性问题仍未解决。

原文摘要 · Abstract (English)

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can establish lower bounds on AI capabilities and assess certain misuse risks given sufficient effort from evaluators. Unfortunately, evaluations face fundamental limitations that cannot be overcome within the current paradigm. These include an inability to establish upper bounds on capabilities, reliably forecast future model capabilities, or robustly assess risks from autonomous AI systems. This means that while evaluations are valuable tools, we should not rely on them as our main way of ensuring AI systems are safe. We conclude with recommendations for incremental improvements to frontier AI safety, while acknowledging these fundamental limitations remain unsolved.

AI安全评估机制风险治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。