揭露大模型评估中的数据污染问题,提出更可靠的评测方法。
A Survey on Data Contamination for Large Language Models
- 分析数据污染成因及对模型性能的虚假提升
- 梳理更新、重写和预防三类无污染评测策略
- 分类对比白盒、灰盒、黑盒检测方法,适合研究者参考
近期大型语言模型(LLMs)在文本生成、代码合成等领域取得显著进展,但其性能评估的可靠性受到数据污染问题的质疑——即训练集与测试集存在意外重叠。由于LLMs通常基于公开来源大规模数据集训练,这些数据集常与评估基准意外重合,导致模型性能被人为高估。本文首先界定数据污染的定义及其影响;其次,回顾无污染评估方法,重点分析基于数据更新、数据重写及预防的三种策略,特别强调动态基准与大模型驱动的评估方法;最后,根据对模型信息的依赖程度,将污染检测方法分为白盒、灰盒和黑盒三类。本综述强调了建立更严格评估协议的必要性,并指明未来应对数据污染挑战的研究方向。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have demonstrated significant progress in various areas, such as text generation and code synthesis. However, the reliability of performance evaluation has come under scrutiny due to data contamination-the unintended overlap between training and test datasets. This overlap has the potential to artificially inflate model performance, as LLMs are typically trained on extensive datasets scraped from publicly available sources. These datasets often inadvertently overlap with the benchmarks used for evaluation, leading to an overestimation of the models' true generalization capabilities. In this paper, we first examine the definition and impacts of data contamination. Secondly, we review methods for contamination-free evaluation, focusing on three strategies: data updating-based methods, data rewriting-based methods, and prevention-based methods. Specifically, we highlight dynamic benchmarks and LLM-driven evaluation methods. Finally, we categorize contamination detecting methods based on model information dependency: white-Box, gray-Box, and black-Box detection approaches. Our survey highlights the requirements for more rigorous evaluation protocols and proposes future directions for addressing data contamination challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。