arXiv:2505.05541cs.AI2025-05综述被引 22

系统梳理AI安全评估方法,帮我们判断模型是否真安全。

Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods

  • 从能力、倾向、抗干扰三方面构建安全评估框架
  • 提出用红队测试、可解释性分析等手段测量极限行为
  • 适合关注AI治理、模型风险控制的研究者和决策者

随着前沿AI系统向变革性能力演进,我们必须同步改进其安全评估方法,以保障系统安全并支持治理决策。尽管基准测试是衡量模型能力的主要方式,但往往无法确立真实上限或预测实际部署表现。本文系统综述了快速发展的AI安全评估领域,提出围绕三个维度的分类体系:测量什么属性、如何测量、以及测量结果如何融入治理框架。评估不仅限于基准,还涵盖模型在极限状态下的能力(如网络攻击、自主复制)、默认行为倾向(如权力追求、策略欺骗),以及面对对抗性AI时安全措施的有效性(控制能力)。这些属性通过行为方法(如支架法、红队测试、监督微调)和内部方法(如表示分析、机制可解释性)进行测量。文章深入解析了若干关键安全风险能力与倾向,并探讨评估结果如何转化为具体研发决策。同时指出评估面临的挑战:证明能力缺失的困难、模型故意隐藏能力(沙袋化)以及‘安全洗白’动机,并指明未来研究方向。本综述旨在整合零散资源,为理解AI安全评估提供核心参考。

原文摘要 · Abstract (English)

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for estimating model capabilities, they often fail to establish true upper bounds or predict deployment behavior. This literature review consolidates the rapidly evolving field of AI safety evaluations, proposing a systematic taxonomy around three dimensions: what properties we measure, how we measure them, and how these measurements integrate into frameworks. We show how evaluations go beyond benchmarks by measuring what models can do when pushed to the limit (capabilities), the behavioral tendencies exhibited by default (propensities), and whether our safety measures remain effective even when faced with subversive adversarial AI (control). These properties are measured through behavioral techniques like scaffolding, red teaming and supervised fine-tuning, alongside internal techniques such as representation analysis and mechanistic interpretability. We provide deeper explanations of some safety-critical capabilities like cybersecurity exploitation, deception, autonomous replication, and situational awareness, alongside concerning propensities like power-seeking and scheming. The review explores how these evaluation methods integrate into governance frameworks to translate results into concrete development decisions. We also highlight challenges to safety evaluations - proving absence of capabilities, potential model sandbagging, and incentives for "safetywashing" - while identifying promising research directions. By synthesizing scattered resources, this literature review aims to provide a central reference point for understanding AI safety evaluations.

AI安全评估方法治理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。