检测到模型分数虚高但排名基本不变,因记忆测试题导致。
Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards
- 用改写题目对比原题,区分记忆与真实能力。
- 实测显示污染仅提升分数,不改变模型排名顺序。
- 建议榜单公布改写题对照结果和置信区间。
基准测试数据泄露(即测试题出现在训练数据中)常被视为大语言模型排行榜可靠性的威胁。本文认为该担忧混淆了两个问题:污染是否虚高绝对得分,以及是否改变模型排名。将污染视为锚点项目不变性破坏,通过原题与语义等价改写题的响应差异来测量,保持能力恒定,分离出记忆效应。在4个基准(ARC、GSM8K、HellaSwag、MMLU)上,对47个公开模型和74个已知污染剂量微调模型的逐项响应进行分析,验证该方法能准确恢复注入的污染程度(测试集泄露导致准确率虚高+0.187点),且对仅使用合法训练集的对照模型无误报(-0.012)。进一步量化排行榜影响:标准榜单与改写控制榜单间排名相关系数达0.997,敏感性分析表明实际污染差异远不足以改变排名,仅有3例模型-基准组合在两份参考中被一致确认存在差异污染。因此,当前公开模型中的污染主要均匀地虚高分数,而不改变排名;排名扭曲需罕见的差异污染。本文提供可校准的不变性审计工具并开放实现,建议排行榜同时报告改写控制排名与置信区间。
原文摘要 · Abstract (English)
Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC, GSM8K, HellaSwag, MMLU), we first calibrate the measure against ground truth: it recovers injected contamination dose-responsively (a corrected effect of +0.187 accuracy points for test-set leakage) and never flags a negative-control model trained only on the legitimate training split (-0.012). We then quantify leaderboard impact: the rank correlation between a standard leaderboard and a paraphrase-controlled leaderboard is 0.997, and a sensitivity analysis shows that the observed differential contamination is far below the level needed to move rankings, with only 3 of 188 model-by-benchmark cases showing differential contamination corroborated across two references. Contamination among these public models is therefore largely uniform: it inflates absolute scores without reordering the leaderboard, and ranking distortion requires the rare case of differential contamination. We provide a calibrated invariance audit, released as a reference implementation, and recommend that leaderboards report paraphrase-controlled rankings alongside confidence intervals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。