arXiv:2606.05241cs.CRcs.AI2026-06被引 4

深度研究型大模型在评测中因网络检索导致性能虚高,需警惕评测污染。

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

论文配图:Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation
图 1 · 摘自论文原文
  • 识别三种检索污染类型:基准元数据泄露、问题上下文泄露、答案直接泄露
  • 实测发现性能虚高最高达4%,现有评测可能严重夸大模型真实推理能力
  • 建议使用隔离沙盒、透明搜索轨迹等措施,提升评测可信度

公开基准测试为大语言模型推理能力的公平与可复现评估提供了可能,但对在推理时主动搜索网络的深度研究型智能体而言变得脆弱。这类智能体可能通过网络检索获取公开基准的元数据、问题上下文甚至真值答案,造成搜索时污染(STC),使外部检索绕过预期推理过程,虚增测评表现。本文系统研究了深度研究代理评估中的STC问题,定义了三种污染程度递增的类型:基准元数据泄露、问题上下文泄露和显式答案泄露,并开发检测算法以识别与量化其影响。在六个公开基准上评估现代深度研究代理,发现STC普遍存在,性能虚高最高可达4%。研究结果表明,现有评估可能高估了模型的真实推理能力。因此,我们倡导采用污染感知实践,包括隔离沙盒、透明搜索轨迹和受控基准访问。

原文摘要 · Abstract (English)

Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such agents may retrieve public benchmark metadata, question context, or even ground-truth answers via web search. This gives rise to Search-Time Contamination (STC), where external retrieval bypasses intended reasoning and inflates measured performance. We systematically study STC in deep research agent evaluation. We define three contamination types with increasing severity, namely Benchmark Metadata Leakage, Question-Context Leakage, and Explicit Answer Leakage, and develop detection algorithms to identify them and quantify their impact on agent performance. Evaluating modern deep research agents on six public benchmarks, we find that STC is widespread and can inflate performance by up to 4%. Our findings show that existing evaluations may overestimate true reasoning ability. We therefore advocate contamination-aware practices, including isolated sandboxes, transparent search trajectories, and controlled benchmark access.

大模型评测推理能力评测污染LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。