arXiv:2605.26438cs.CLcs.AI2026-05

提出LURE方法,让大模型在更真实的交互中被评估,减少因感知被评测而表现失真。

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

论文配图:LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
图 1 · 摘自论文原文
  • 通过重放真实智能体交互轨迹并末尾添加评测提示,模拟部署场景
  • 实验表明其评测结果与真实用户对话接近,显著优于现有基准
  • 适合用于安全测试、对齐评估等需高真实性的场景

大语言模型能识别自身正在被评估(评估意识),从而改变行为,削弱了安全与对齐基准的有效性。我们提出LURE(实时使用回放评估),通过重放真实智能体交互轨迹并在末尾添加评测提示,构建类部署评估环境。我们还设计了自动化管道来衡量评估真实性,结合显式表达评估意识的检测与判断模型对日志是否为评测的概率估计,并在大规模部署与评测语料库上验证。结果显示,基于LURE的评估与真实部署场景的可区分度显著低于主流基准和合成生成器,可逼近真实用户对话的真实感。我们在诡计、AI安全破坏和奉承行为等场景中实现了LURE实例化。结果表明,评估真实性是对齐基准的关键属性,应在报告基准结果时一并披露,尤其当结果用于安全论证时。

原文摘要 · Abstract (English)

Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignment benchmarks. We propose LURE (Live-Usage Replay Evaluations), a method for constructing deployment-like evaluations by replaying realistic agentic interaction trajectories and appending evaluation prompt at the end. We also introduce an automated pipeline for measuring evaluation realism, combining detection of verbalized evaluation awareness and judge-model estimates of the probability of logs being an evaluation, and validate it on a large dataset of deployment and evaluation transcripts. We find that LURE-based evaluations are substantially less distinguishable from deployment than widely used benchmarks and synthetic evaluation generators, and can approach the realism of real conversations with users. We instantiate LURE in scheming, AI safety sabotage, and sycophancy settings. Our results suggest that evaluation realism is a crucial property of alignment benchmarks and should be reported alongside benchmark results, especially when such results are used in safety cases.

评估真实度对齐测试LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。