arXiv:2602.20813cs.AI2026-02

通过904个真实压力场景评估大模型对齐性,发现顶级模型仍有短板。

Pressure Reveals Character: Behavioural Alignment Evaluation at Depth

  • 设计6类共904个多轮对抗场景,模拟真实压力下的行为表现
  • 24个前沿模型测试中,多数在多个类别存在持续弱点
  • 发现对齐能力呈统一特质,高分模型在各维度普遍表现好

评估语言模型的对齐性需在真实压力下检验其行为,而非仅依赖自我声明。随着对齐失败造成越来越多现实危害,具备真实多轮场景的综合评估框架仍匮乏。我们构建了一个涵盖6个类别(诚实、安全、非操控、鲁棒性、可纠正性、阴谋性)共904个场景的对齐基准,经人类评分验证其真实性。这些场景引入冲突指令、模拟工具访问与多轮升级,揭示单轮评估无法捕捉的行为倾向。使用基于大模型的裁判评估24个前沿模型,结果表明即使顶尖模型在特定类别仍存在差距,而多数模型在整体上表现出一致弱点。因子分析显示,对齐能力呈现统一结构(类似认知研究中的g因子),在某一类别得分高的模型往往在其他类别也表现优异。我们已公开该基准与交互式排行榜,后续将扩展薄弱领域场景,并持续纳入新发布模型。

原文摘要 · Abstract (English)

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cause real-world harm, comprehensive evaluation frameworks with realistic multi-turn scenarios remain lacking. We introduce an alignment benchmark spanning 904 scenarios across six categories -- Honesty, Safety, Non-Manipulation, Robustness, Corrigibility, and Scheming -- validated as realistic by human raters. Our scenarios place models under conflicting instructions, simulated tool access, and multi-turn escalation to reveal behavioural tendencies that single-turn evaluations miss. Evaluating 24 frontier models using LLM judges validated against human annotations, we find that even top-performing models exhibit gaps in specific categories, while the majority of models show consistent weaknesses across the board. Factor analysis reveals that alignment behaves as a unified construct (analogous to the g-factor in cognitive research) with models scoring high on one category tending to score high on others. We publicly release the benchmark and an interactive leaderboard to support ongoing evaluation, with plans to expand scenarios in areas where we observe persistent weaknesses and to add new models as they are released.

模型对齐行为评估大模型测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。