arXiv:2606.14965cs.LGcs.DB2026-06

通过可控数据扰动构建真实可解释的标签噪声基准,揭示算法在不同噪声下的隐藏缺陷。

Benchmarking Instance-Dependent Label Noise with Controlled Corruptions

论文配图:Benchmarking Instance-Dependent Label Noise with Controlled Corruptions
图 1 · 摘自论文原文
  • 用可控输入扰动生成实例相关噪声,明确噪声来源与严重程度。
  • 在CIFAR-10等数据集上构建90种噪声场景,噪声分布更贴近人类不确定性。
  • 发现现有方法在扰动噪声下暴露新失败模式,证明噪声结构影响算法表现。

合成实例相关标签噪声(IDN)基准广泛用于评估噪声标签学习方法,但现有方法通常通过不完美标注者或分类器评分生成噪声,模糊了歧义来源。本文提出CILN框架,通过可控输入扰动生成IDN。多样化的标注者群体对扰动样本进行标注,生成具有明确且可调控的歧义源与严重度的基准数据集。基于CIFAR10、MNIST和Adult数据集,构建了90个涵盖多种扰动类型与严重级别的基准设置。实验表明,该基准产生真实的实例相关噪声,具备多样的混淆结构;在CIFAR-10上,其标签分布比现有合成基准更接近人类不确定性。此外,基于扰动的IDN能揭示主流方法(如Co-Teaching和DivideMix)在同类误标率下未被察觉的失败模式。结果表明,噪声结构不仅噪声率,对基准难度与算法行为均有关键影响。通过显式可控的歧义生成,CILN为研究多样化实例困难下的噪声标签学习提供了互补性基准框架。

原文摘要 · Abstract (English)

Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit. We introduce CILN, a benchmark generation framework that creates IDN through controlled input corruptions. A diverse voter pool labels corrupted instances, producing benchmark datasets in which both the source and severity of ambiguity are explicit and controllable. Using CIFAR10, MNIST, and Adult, we construct 90 benchmark settings spanning multiple corruption families and severity levels. Our experiments show that the resulting benchmarks exhibit genuine instance-dependent noise, provide diverse confusion structures, and, on CIFAR-10, can produce label distributions that are closer to human uncertainty than an existing synthetic IDN benchmark. We further demonstrate that corruption-mediated IDN can expose failure modes of popular noisy-label learning methods, including Co-Teaching and DivideMix, that are not observed under comparable levels of rater-fallibility noise. These findings suggest that noise structure, not only noise rate, plays an important role in benchmark difficulty and algorithm behavior. By making ambiguity generation explicit and controllable, CILN provides a complementary benchmarking framework for studying noisy-label learning under diverse sources of instance difficulty.

噪声学习基准测试数据扰动实例依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。