实证分析2047个数据集,发现选择性泄露比归一化泄露更致命。
Which Leakage Types Matter? A Quantitative Landscape Across 2,047 Benchmark Datasets

- 通过28组对照实验,量化四类数据泄露的严重程度
- 选择性泄露导致报告得分虚高约90%噪声利用效应
- 模型容量越大,记忆泄露越严重,适合模型评估者阅读
在2,047个独立同分布的表格数据集上开展28组受试者内反事实实验,并在129个时序数据集上进行边界实验,量化机器学习中四类数据泄露的严重程度。第一类(估计:用全量数据拟合归一化器)可忽略不计:所有九种情况的|ΔAUC|均≤0.005。第二类(选择:窥探、种子挑拣)影响显著:测量效应相当于约90%噪声被利用而夸大得分。第三类(记忆)随模型容量增加:在10%数据重复下,朴素贝叶斯为d_z=0.37,决策树达d_z=1.11。第四类(边界)在随机交叉验证下不可见。在此独立同分布表格数据场景中,教科书强调顺序被颠倒:归一化泄露影响最小;在实际数据规模下,选择泄露影响最大。
原文摘要 · Abstract (English)
Twenty-eight within-subject counterfactual experiments across 2,047 iid tabular datasets, plus a boundary experiment on 129 temporal datasets, measure the severity of four data leakage classes in machine learning. Class I (estimation: fitting scalers on full data) is negligible: all nine conditions produce $|ΔAUC| \leq 0.005$. Class II (selection: peeking, seed cherry-picking) is substantial: the measured effect is consistent with about 90% noise exploitation inflating reported scores. Class III (memorization) scales with model capacity: $d_z$ = 0.37 (Naive Bayes) to 1.11 (Decision Tree) at 10% duplication. Class IV (boundary) is invisible under random cross-validation. Within this iid tabular regime, the textbook emphasis is inverted: normalization leakage matters least; selection leakage at practical dataset sizes matters most.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。