arXiv:2606.30219cs.AIcs.CL2026-06综述

揭示大模型评估与安全之间的测量鸿沟,提出新框架统一理解评价偏差。

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

论文配图:EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
图 1 · 摘自论文原文
  • 构建八维度证据分类体系,整合373篇研究
  • 发现评估分数提升不等于能力或对齐真正改善
  • 适合关注模型安全评估的开发者和审计者

本文系统综述并融合了2018至2026年间373篇主要研究,揭示大语言模型(LLM)评估与人工智能安全中的共性问题:基准分数、奖励信号与安全度量虽在提升,但其所代表的能力与对齐属性仍不确定。研究将基准有效性、污染问题、动态评估、LLM作为裁判协议、对抗性安全测试、奖励与代理优化、机制可解释性及人工智能治理等证据归纳为八流证据分类体系。在此基础上,提出EvalSafetyGap概念框架,将基准有效性与对齐失败统一为优化压力下的代理-目标偏差问题,通过受古德哈特启发的不稳定性分解与对齐三难困境形式化表达。十模型公开证据审计表明,应将能力、行为鲁棒性和治理披露作为独立证据层呈现,而非合并为单一安全评分。最后提出动态抗污染基准、预设多轮威胁模型、版本锁定评估、透明来源报告及验证的机制安全指标等研究议程,为研究人员、模型开发者与AI审计者提供测量意识导向的安全评估共同语言。

原文摘要 · Abstract (English)

This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.

大模型评估安全对齐测量偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。