arXiv:2411.03923cs.CL2024-11被引 35

提出新方法评估大模型评测数据污染,揭示其对性能影响被低估

Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

  • 通过下游任务效果反推污染样本,建立可验证的污染评估框架
  • 发现污染影响远超现有报告,且在不同模型规模下作用各异
  • 建议按模型和基准定制阈值,避免误判,提升分析精度

评测数据污染严重干扰大模型性能评估结果的解读,但如何定义污染样本及其影响仍不明确。本文提出一种新分析方法ConTAM,将污染判断与模型实际收益关联,通过在13个基准、7个来自两个不同家族的模型上测试多种n-gram污染指标,发现污染的影响远大于近期模型发布所报告的水平,且在不同模型规模下表现不同。研究还表明,仅考虑最长污染子串比合并所有污染子串更具信号性;针对模型和基准分别设置阈值可显著提高结果特异性。此外,实验显示增大n值或忽略预训练数据中不频繁的匹配会引入大量假阴性。ConTAM为污染度量提供了基于下游效应的实证基础,揭示了污染对大模型的真实影响,并为未来分析提供具体建议。

原文摘要 · Abstract (English)

Hampering the interpretation of benchmark scores, evaluation data contamination has become a growing concern in the evaluation of LLMs, and an active area of research studies its effects. While evaluation data contamination is easily understood intuitively, it is surprisingly difficult to define precisely which samples should be considered contaminated and, consequently, how it impacts benchmark scores. We propose that these questions should be addressed together and that contamination metrics can be assessed based on whether models benefit from the examples they mark contaminated. We propose a novel analysis method called ConTAM, and show with a large scale survey of existing and novel n-gram based contamination metrics across 13 benchmarks and 7 models from 2 different families that ConTAM can be used to better understand evaluation data contamination and its effects. We find that contamination may have a much larger effect than reported in recent LLM releases and benefits models differently at different scales. We also find that considering only the longest contaminated substring provides a better signal than considering a union of all contaminated substrings, and that doing model and benchmark specific threshold analysis greatly increases the specificity of the results. Lastly, we investigate the impact of hyperparameter choices, finding that, among other things, both using larger values of n and disregarding matches that are infrequent in the pre-training data lead to many false negatives. With ConTAM, we provide a method to empirically ground evaluation data contamination metrics in downstream effects. With our exploration, we shed light on how evaluation data contamination can impact LLMs and provide insight into the considerations important when doing contamination analysis. We end our paper by discussing these in more detail and providing concrete suggestions for future work.

大模型评估数据污染评测可信度ConTAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。