arXiv:2605.24818stat.MEcs.CL2026-05

通过在训练集注入已知污染样本,实现对测试集分数的精准修正。

Correcting test set contamination by spiking the training data

论文配图:Correcting test set contamination by spiking the training data
图 1 · 摘自论文原文
  • 在训练数据中故意掺入测试样本,构建可校准的记忆化预测器。
  • 使用记忆与正确性双重信息的修正方法比无修正更准确。
  • 仅需10个样本即可校准简单预测器,且跨数据集通用。

现有研究多关注测试集污染的检测,而对污染后测试分数的修正研究不足。本文提出在训练数据中以已知比例故意掺入部分测试样例,利用这些被标记的污染样本校准模型记忆程度的预测器,从而实现对虚高测试分数的统计修正。为评估不同修正估计器,我们基于哈勃模型构建了仿真框架:哈勃模型成对出现,其中扰动模型被故意污染多个测试集,标准模型则未被污染,作为反事实对照和修正目标。我们考察了使用记忆预测、正确性预测或两者结合的估计器,在模拟实验中验证了结合记忆与正确性信息的方法优于不作修正的基线方法。进一步实例化多种记忆与正确性预测器,发现普拉特缩放的记忆推理指标等简单方法已具备良好修正信号。最后分析了注入实践的可行性:简单记忆预测器仅需约10个样本即可校准,且常可在不同数据集间迁移。总体表明,注入法是应对测试集污染的有力方案。

原文摘要 · Abstract (English)

The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core proposal is to spike the training data by intentionally contaminating some test examples at known rates. The spiked examples can then be used to calibrate predictors of model memorization which enable principled statistical correction of inflated test scores. To evaluate different correction estimators, we first present a simulation framework based on the Hubble models. Hubble models come in minimal pairs, where the perturbed model was deliberately contaminated with several test sets, while the standard model was not, serving as the counterfactual and correction target. We consider estimators that use information from a memorization predictor, correctness predictor, or both. In simulation, we establish basic statistical intuitions and show that estimators leveraging memorization and correctness information are better than naive estimation which makes no correction at all. We then instantiate several memorization and correctness predictors, and find that simple predictors such as Platt-scaled membership inference metrics provide good signal for correction. Finally, we examine the practical considerations of spiking. Simple memorization predictors need no more than 10 examples for calibration and often transfer from one dataset to another. Taken together, spiking is a promising solution for test set contamination.

测试集污染模型记忆统计修正数据注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。