arXiv:2505.01420cs.LG2025-05被引 28

测试大模型是否具备隐蔽行动和环境感知能力,评估其潜在失控风险。

Evaluating Frontier Models for Stealth and Situational Awareness

  • 设计五项隐蔽性推理测试,评估模型规避监督的能力
  • 构建十一个情境意识测试,衡量模型自省与环境建模水平
  • 当前前沿模型在两项能力上均未表现出危险水平

近期研究揭示了前沿人工智能模型可能暗中策划——明知且隐蔽地追求与开发者意图不符的目标。此类行为极难检测,若出现在未来高级系统中,可能带来严重的控制权丧失风险。因此,在部署前排除模型的谋划行为至关重要。本文提出一套用于评估谋划推理能力的测评体系,包含两类关键能力:一是五项关于规避监管的隐蔽性推理测试;二是十一项衡量模型对自身、环境及部署状态进行工具性推理的情境意识测试。我们证明这些测评可构成‘无谋划能力’的安全论证基础:若模型未能通过这些测试,则几乎不可能在真实部署中造成严重危害。我们在当前前沿模型上进行了测试,发现它们在情境意识和隐蔽性方面均未表现出令人担忧的水平。

原文摘要 · Abstract (English)

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future advanced systems, could pose severe loss of control risk. It is therefore important for AI developers to rule out harm from scheming prior to model deployment. In this paper, we present a suite of scheming reasoning evaluations measuring two types of reasoning capabilities that we believe are prerequisites for successful scheming: First, we propose five evaluations of ability to reason about and circumvent oversight (stealth). Second, we present eleven evaluations for measuring a model's ability to instrumentally reason about itself, its environment and its deployment (situational awareness). We demonstrate how these evaluations can be used as part of a scheming inability safety case: a model that does not succeed on these evaluations is almost certainly incapable of causing severe harm via scheming in real deployment. We run our evaluations on current frontier models and find that none of them show concerning levels of either situational awareness or stealth.

AI安全大模型风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。