arXiv:2605.21543cs.LG2026-05

提出可证明的联合去污染方法,确保多模型评估公平性

Provable Joint Decontamination for Benchmarking Multiple Large Language Models

  • 联合建模多个模型的训练数据污染,统一控制全局污染率
  • 在多种模型和基准上实现更高检测力且严格满足目标污染率
  • 适合关注大模型评测可信度的研究者与实践者

基准数据污染已成为大语言模型评估的核心挑战:当评测样本出现在一个或多个被测模型的训练数据中时,报告性能可能被夸大,跨模型比较变得不可靠。现有基于打分的训练数据检测方法缺乏理论保障;虽有近期的同质化方法为单个模型提供可证明的误识别控制,但对每个模型分别应用会生成模型专属基准,破坏公平比较。本文将多模型基准去污染形式化为联合选择问题,提出联合包络同质选择(JECS)方法,在给定假设下实现全局污染率(GCR)的可证明控制。具体而言,JECS计算各模型的同质p值,按项取最大值进行聚合,并从右尾观测中重构最大p值的零分布保守包络。通过将自适应Benjamini-Hochberg(BH)程序应用于包络重缩放后的值,选取具有可证明GCR控制的基准。在多种模型与基准上的大量实验表明,JECS在保持目标GCR控制的同时,检测效能显著优于最大p值基线。

原文摘要 · Abstract (English)

Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performance can be inflated and cross-model comparisons become unreliable. A broad line of training-data detection work designs scores to quantify how strongly a model memorizes a given data point, but these score-based methods lack theoretical guarantees. Recent conformal approaches provide provable false-identification control for a single model; however, applying them separately to each model can produce model-specific benchmarks, undermining fair comparison across models. In this work, we formalize multi-model benchmark decontamination as a joint selection problem and propose Joint Envelope Conformal Selection (JECS), a conformal procedure that enables global contamination rate (GCR) control under stated assumptions. Specifically, JECS computes per-model conformal p-values, aggregates them by the per-item maximum, and reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold. By applying the adaptive Benjamini-Hochberg (BH) procedure to the envelope-rescaled values, we select a benchmark with provable GCR control. Extensive experiments across various models and benchmarks demonstrate that JECS achieves higher power than the max-p baseline while consistently maintaining the target GCR control.

大模型评测数据污染统计推断可证明性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。