提出新框架检测表格数据中的模型过拟合,发现多个主流数据集存在严重污染。
When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets
- 设计可控查询生成与对比评估,精准区分模型泛化与数据泄露。
- 在8个数据集中发现4个存在显著污染,性能虚高风险大。
- 适合关注数据可信度、评估严谨性的机器学习研究者使用。
大型语言模型在表格数据上的表现可能受数据污染影响,即因提前接触测试集而获得虚假性能提升,而非真正泛化能力。现有方法多依赖粗粒度的记忆测试,难以有效识别污染。本文提出一种新框架:通过构造结构一致的多选题查询,并系统性地变换数据内容,实现对信息泄露的精准探测。这些变换可选择性破坏原始数据,同时保留部分知识,从而分离出由污染带来的性能增益。研究还引入非神经基线和统计检验,以判断性能异常是否显著。在8个常用表格数据集上的实证结果表明,其中4个存在明显污染证据,暗示当前下游任务评估可能存在严重性能虚高,对评估可靠性构成挑战。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this problem is largely unexplored. Existing approaches primarily rely on memorization tests, which are too coarse to detect contamination. In contrast, we propose a framework for assessing contamination in tabular datasets by generating controlled queries and performing comparative evaluation. Given a dataset, we craft multiple-choice aligned queries that preserve task structure while allowing systematic transformations of the underlying data. These transformations are designed to selectively disrupt dataset information while preserving partial knowledge, enabling us to isolate performance attributable to contamination. We complement this setup with non-neural baselines that provide reference performance, and we introduce a statistical testing procedure to formally detect significant deviations indicative of contamination. Empirical results on eight widely used tabular datasets reveal clear evidence of contamination in four cases. These findings suggest that performance on downstream tasks involving such datasets may be substantially inflated, raising concerns about the reliability of current evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。