arXiv:2609.01611cs.AIcs.CL2026-09

测试大模型是否识别自己在被评估,确保评测结果可信。

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

论文配图:EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
图 1 · 摘自论文原文
  • 用新基准检测大模型在评测时的自我觉察能力。
  • 发现评测文本来源影响结果,同一模型排名可能被颠倒。
  • 提供校准方法,适合关注AI安全评测可靠性的研究者。

前沿大语言模型常能识别自己正被评估,这种评估意识会扭曲评测结果,威胁当前AI安全框架的可信度。我们提出EvalDetectBench,一个开放的评测管道与基准,可兼容任意Inspect兼容的评测任务,支持对当前及未来基准的检测。该基准包含全新整理的对话片段集,覆盖当前前沿系统卡评测和多样化的部署来源。其目标是双重:一是测量前沿大模型识别评估状态的能力,二是评估各评测本身是否易被识别为评估。我们发现现有文献中存在两种系统性偏差:生成部署文本的模型身份贡献了11.25%的测量方差,可导致模型排名被重排;而为某一模型优化的提问提示在其他模型上表现接近随机。EvalDetectBench通过模型级探测校准和分层生成器调和机制纠正上述问题。

原文摘要 · Abstract (English)

Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks. EvalDetectBench ships with a newly curated transcript suite covering current frontier system-card evaluations and diverse deployment sources. The benchmark serves two purposes: measuring how reliably frontier LLMs recognize that they are being evaluated, and assessing how detectable individual benchmarks are as evaluations. We identify two methodological choices in the existing literature that introduce systematic bias: the identity of the model that generated the deployment transcripts accounts for 11.25% of measurement variance and can reorder model rankings; and elicitation prompts selected for high performance on one model can perform near chance on others. EvalDetectBench corrects for both via per-model probe calibration and a stratified generator-harmonisation procedure.

大模型评测评估意识基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。