提出新方法提升模型在噪声标签和域偏移下的泛化能力。
Noise-Aware Generalization: Robustness to In-Domain Noise and Out-of-Domain Generalization
- 利用跨域样本差异识别噪声,避免误将域偏移当作噪声。
- 在7个数据集上最高提升12.5%性能,优于现有方法组合。
- 适合需要鲁棒训练的工业场景和多源数据任务。
针对标签噪声与域偏移并存的挑战,本文提出首个直接解决噪声感知泛化的方法DL4ND。传统域泛化(DG)方法在标签噪声下效果显著下降,而学习噪声标签(LNL)方法则易对易学域过拟合,误将域变化视为噪声。本研究发现,单域内难以区分的噪声样本,在跨域对比中表现出更大差异,据此设计了基于域标签的噪声检测机制。实验表明,即使使用域标签分离噪声与域偏移,DL4ND仍显著优于现有DG与LNL方法及其组合,在7个不同数据集、3种噪声类型下,性能提升最高达12.5%。
原文摘要 · Abstract (English)
Methods addressing Learning with Noisy Labels (LNL) and multi-source Domain Generalization (DG) use training techniques to improve downstream task performance in the presence of label noise or domain shifts, respectively. Prior work often explores these tasks in isolation, and the limited work that does investigate their intersection, which we refer to as Noise-Aware Generalization (NAG), only benchmarks existing methods without also proposing an approach to reduce its effect. We find that this is likely due, in part, to the new challenges that arise when exploring NAG, which does not appear in LNL or DG alone. For example, we show that the effectiveness of DG methods is compromised in the presence of label noise, making them largely ineffective. Similarly, LNL methods often overfit to easy-to-learn domains as they confuse domain shifts for label noise. Instead, we propose Domain Labels for Noise Detection (DL4ND), the first direct method developed for NAG which uses our observation that noisy samples that may appear indistinguishable within a single domain often show greater variation when compared across domains. We find DL4ND outperforms DG and LNL methods, including their combinations, even when simplifying the NAG challenge by using domain labels to isolate domain shifts from noise. Performance gains up to 12.5% over seven diverse datasets with three noise types demonstrates DL4ND's ability to generalize to a wide variety of settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。