arXiv:2510.02386cs.CRcs.AI2025-10中稿 · ICLR被引 6

推理模型评估易被污染,现有检测方法几乎失效。

On The Fragility of Benchmark Contamination Detection in Reasoning Models

  • 微调后用强化学习快速隐藏数据污染痕迹
  • 多数检测方法在新模型上准确率接近随机
  • 适用于关注评估公平性的研究者与评测平台

推理模型(LRM)的排行榜促使开发者直接优化于评测集,导致将评测数据融入训练以虚增性能,即评测污染。我们发现,即使在微调阶段已可检测污染,仅经过短暂的GRPO训练,多数检测方法依赖的关键信号便会被显著掩盖。实证与理论分析表明,PPO类重要性采样和裁剪目标是造成检测失效的根本原因,暗示大量强化学习方法可能具有相似隐藏能力。此外,当使用思维链(CoT)进行最终微调时,大多数检测方法表现近乎随机。由于污染模型对分布相似但未见样本仍保持高置信度,使其能规避基于记忆的检测。综上,当前评测体系存在严重漏洞:开发者可轻易污染模型并留下极小痕迹,严重威胁排行榜公正性与可信度。亟需针对推理模型设计更先进的检测方法与可信评估协议。

原文摘要 · Abstract (English)

Leaderboards for LRMs have turned evaluation into a competition, incentivizing developers to optimize directly on benchmark suites. A shortcut to achieving higher rankings is to incorporate evaluation benchmarks into the training data, thereby yielding inflated performance, known as benchmark contamination. Surprisingly, our studies find that evading contamination detections for LRMs is alarmingly easy. We focus on the two scenarios where contamination may occur in practice: (I) when the base model evolves into LRM via SFT and RL, we find that contamination during SFT can be originally identified by contamination detection methods. Yet, even a brief GRPO training can markedly conceal contamination signals that most detection methods rely on. Further empirical experiments and theoretical analysis indicate that PPO style importance sampling and clipping objectives are the root cause of this detection concealment, indicating that a broad class of RL methods may inherently exhibit similar concealment capability; (II) when SFT contamination with CoT is applied to advanced LRMs as the final stage, most contamination detection methods perform near random guesses. Without exposure to non-members, contaminated LRMs would still have more confidence when responding to those unseen samples that share similar distributions to the training set, and thus, evade existing memorization-based detection methods. Together, our findings reveal the unique vulnerability of LRMs evaluations: Model developers could easily contaminate LRMs to achieve inflated leaderboards performance while leaving minimal traces of contamination, thereby strongly undermining the fairness of evaluation and threatening the integrity of public leaderboards. This underscores the urgent need for advanced contamination detection methods and trustworthy evaluation protocols tailored to LRMs.

模型评估数据污染强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。