arXiv:2608.06511cs.LG2026-08

提出匹配操作点评估框架,解决自适应数据清洗中的移除预算混淆问题。

Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning

  • 设计匹配移除预算与召回率的评估框架,消除因分组粒度变化导致的偏差。
  • 在CIFAR-10和ImageNet-100上验证,多数性能提升在统一操作点后消失。
  • 发现难样本是低污染率下错误的主要来源,适合关注低误报场景的研究者参考。

自适应数据清洗方法用数据驱动的分组替代人工阈值。但改变分组粒度(即按估计污染风险划分样本的组数)会隐式移动决策边界,影响整体移除样本数,造成移除预算混淆——看似精度或假阳性率提升,实则源于更小的移除预算。为此,我们提出一种操作点感知的评估框架,通过匹配预算与召回率控制,结合阈值无关指标(AUROC、AUPRC)进行评估。在包含重加权学习难度线索、辅助欧氏距离线索及更高分组粒度的多线索清洗器重构实验中,朴素评估显示显著性能提升,但在操作点匹配后这些增益消失。假阳性分解表明:在低污染率下,干净但困难样本主导错误;中等污染时变为阈值依赖;严重污染下贡献可忽略。在CIFAR-10和ImageNet-100上的实验显示,大多数性能差异在低至中等污染水平下经操作点匹配后缩小或消失。真正的排名优势仅出现在特定低频场景及严重污染下的高召回区域。结果强调,自适应清洗方法必须在匹配操作点下评估,才能确保性能提升反映真实污染识别能力。

原文摘要 · Abstract (English)

Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics like precision or false-positive rate reflect a smaller removal budget rather than superior corruption discrimination. To address this evaluation bias, we introduce an operating-point-aware evaluation framework that evaluates methods using matched-budget and matched-recall controls alongside threshold-independent metrics (AUROC and AUPRC). We test this framework on a multi-cue adaptive cleaner redesign featuring a reweighted learning-difficulty cue, an auxiliary Euclidean-distance cue, and increased partition granularity intended to isolate clean-but-difficult samples. While naive evaluations (assessing configurations at their own induced operating points) suggest substantial performance improvements for the redesign, these gains disappear once operating points are equalized. False-positive decomposition reveals that clean-but-difficult samples primarily drive error counts at low corruption rates, become threshold-dependent at moderate corruption, and contribute negligibly under severe corruption. Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched. True ranking advantages only remain in specific low-prevalence settings and in high-recall regions under severe corruption. These findings highlight that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.

数据清洗评估框架机器学习模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。