arXiv:2503.11414cs.LG2025-03被引 4

针对长尾与标签噪声共存问题,提出解耦去噪新方法提升模型鲁棒性。

Classifying Long-tailed and Label-noise Data via Disentangling and Unlearning

  • 通过特征解耦分离内部表示,识别错误关联特征
  • 引入部分去学习机制削弱错误类别特征影响,缓解尾部样本误标为头部现象
  • 适用于真实长尾带噪数据,尤其适合标注质量差的场景

现实数据中长尾分布与标签噪声常同时存在,严重阻碍模型训练。现有研究通常假设噪声生成与长尾无关,但实际中尾部类样本更易被误标为头部类,加剧不平衡问题,我们称之为“尾到头(T2H)”噪声。该噪声污染头部类,迫使模型将尾部样本误认为头部,显著降低性能。为此,本文提出面向长尾与标签噪声数据的解耦去学习方法(DULL)。首先利用内特征解耦(IFD)分离内部特征表示,再通过内特征部分去学习(IFPU)弱化与错误类别相关联的特征区域,防止模型被噪声误导,增强对噪声的鲁棒性。为构建可控实验环境,还设计了一种模拟T2H噪声的新算法。在模拟和真实数据集上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

In real-world datasets, the challenges of long-tailed distributions and noisy labels often coexist, posing obstacles to the model training and performance. Existing studies on long-tailed noisy label learning (LTNLL) typically assume that the generation of noisy labels is independent of the long-tailed distribution, which may not be true from a practical perspective. In real-world situaiton, we observe that the tail class samples are more likely to be mislabeled as head, exacerbating the original degree of imbalance. We call this phenomenon as ``tail-to-head (T2H)'' noise. T2H noise severely degrades model performance by polluting the head classes and forcing the model to learn the tail samples as head. To address this challenge, we investigate the dynamic misleading process of the nosiy labels and propose a novel method called Disentangling and Unlearning for Long-tailed and Label-noisy data (DULL). It first employs the Inner-Feature Disentangling (IFD) to disentangle feature internally. Based on this, the Inner-Feature Partial Unlearning (IFPU) is then applied to weaken and unlearn incorrect feature regions correlated to wrong classes. This method prevents the model from being misled by noisy labels, enhancing the model's robustness against noise. To provide a controlled experimental environment, we further propose a new noise addition algorithm to simulate T2H noise. Extensive experiments on both simulated and real-world datasets demonstrate the effectiveness of our proposed method.

长尾学习标签噪声去学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。