利用激活值大小提升噪声数据下的分类鲁棒性
Robust Partial-Label Learning by Leveraging Class Activation Values
- 基于主观逻辑,用激活值大小表示不确定性
- 在高噪声、分布外和对抗样本下表现更稳定
- 适合处理标注混乱的真实数据,提升模型可靠性
真实训练数据常含噪声,例如人工标注对同一实例给出冲突标签。部分标签学习(PLL)是一种弱监督学习范式,可在无需手动清理数据的情况下训练分类器。尽管当前先进方法预测性能良好,但其预测对高噪声、分布外数据及对抗扰动仍敏感。本文提出一种基于主观逻辑的新型PLL方法,通过利用神经网络类别激活值的大小显式表示不确定性,并设计了一种证明为最优的标签权重重分配策略,有效融入先验标签知识。实验表明,该方法在高PLL噪声水平、分布外样本以及测试实例的对抗扰动下均能实现更稳健的预测性能。
原文摘要 · Abstract (English)
Real-world training data is often noisy; for example, human annotators assign conflicting class labels to the same instances. Partial-label learning (PLL) is a weakly supervised learning paradigm that allows training classifiers in this context without manual data cleaning. While state-of-the-art methods have good predictive performance, their predictions are sensitive to high noise levels, out-of-distribution data, and adversarial perturbations. We propose a novel PLL method based on subjective logic, which explicitly represents uncertainty by leveraging the magnitudes of the underlying neural network's class activation values. Thereby, we effectively incorporate prior knowledge about the class labels by using a novel label weight re-distribution strategy that we prove to be optimal. We empirically show that our method yields more robust predictions in terms of predictive performance under high PLL noise levels, handling out-of-distribution examples, and handling adversarial perturbations on the test instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。