提出抗污染基准数据集,确保大模型评估更可靠
LLM Benchmark Datasets Should Be Contamination-Resistant

- 利用Transformer架构的推理与训练差异实现数据不可学但可推理
- 发现多数基准数据已污染,削弱模型泛化评估效力
- 呼吁构建抗污染评估体系,适合模型评测研究者参考
基准数据集对大语言模型的可复现、可靠与区分性评估至关重要。然而近期研究发现,许多基准数据集已被包含在预训练语料中,即存在污染问题,从而削弱其作为模型泛化能力衡量标准的价值。本文主张基准数据集应具备抗污染性,即不可被模型学习,但仍支持推理。为此,我们首先揭示了基准数据集污染的广泛存在,并定义了抗污染数据集的特性;其次,指出Transformer架构中推理与训练路径的不对称性可被用于实现抗污染性;最后,提出数学方法以实现跨不同大模型架构的数据集互操作。基于此,我们呼吁社区通过三方面举措提升评估可靠性:(i)发展新型抗污染方法,(ii)构建支持平台,(iii)将抗污染基准纳入现有评估流程。
原文摘要 · Abstract (English)
Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark datasets are included in pretraining corpora, i.e., $\textit{contaminated}$, which diminishes their value as reliable measures of model generalization. In this paper, we argue that benchmark datasets should be $\textit{contamination-resistant}$, i.e., $\textit{unlearnable}$, but support $\textit{inference}$. To accomplish this, we first highlight the wide prevalence of benchmark dataset contamination and outline the properties of contamination-resistant datasets. Second, we highlight how the asymmetry between the inference and training pipelines in the Transformer architecture can be leveraged to support contamination-resistance. Third, we outline mathematical advancements to make these datasets interoperable across various LLM architectures. Based on the above, we call on the community to ensure the reliability of LLM benchmarking by: (i) advancing novel contamination-resistant methodologies, (ii) developing supporting methods and platforms, and (iii) adopting contamination-resistant benchmarks into existing evaluation pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。