arXiv:2502.14281cs.LGcs.AI2025-02被引 2

通过潜在空间偏移建模标签噪声,提升多标签预测准确性

Correcting Noisy Multilabel Predictions: Modeling Label Noise through Latent Space Shifts

  • 将标签噪声视为潜在变量的随机偏移,构建生成模型
  • 在多种噪声设置下均优于现有方法,显著提升预测性能
  • 适用于已有模型的后处理,适合实际应用中噪声修正

数据噪声在多数真实机器学习场景中不可避免,常导致严重过拟合。不仅特征可能含噪,标签也因人工标注易出错。本文聚焦于少被研究的多标签分类中的标签噪声问题,重点研究基于已训练模型生成预测后的后修正方法。该策略可直接复用已有模型以节省计算资源,并能与其他噪声修正技术结合进一步提升效果。我们采用深度生成模型进行不确定性估计,假设标签噪声源于潜在变量的随机变化,从而提供更鲁棒的噪声学习机制。提出无监督与半监督两种学习方法,实验证明其在各类噪声条件下均能持续改进独立模型性能,优于多个现有方法。通过敏感性分析、消融实验等全面评估,验证了方法的稳健性。

原文摘要 · Abstract (English)

Noise in data appears to be inevitable in most real-world machine learning applications and would cause severe overfitting problems. Not only can data features contain noise, but labels are also prone to be noisy due to human input. In this paper, rather than noisy label learning in multiclass classifications, we instead focus on the less explored area of noisy label learning for multilabel classifications. Specifically, we investigate the post-correction of predictions generated from classifiers learned with noisy labels. The reasons are two-fold. Firstly, this approach can directly work with the trained models to save computational resources. Secondly, it could be applied on top of other noisy label correction techniques to achieve further improvements. To handle this problem, we appeal to deep generative approaches that are possible for uncertainty estimation. Our model posits that label noise arises from a stochastic shift in the latent variable, providing a more robust and beneficial means for noisy learning. We develop both unsupervised and semi-supervised learning methods for our model. The extensive empirical study presents solid evidence to that our approach is able to consistently improve the independent models and performs better than a number of existing methods across various noisy label settings. Moreover, a comprehensive empirical analysis of the proposed method is carried out to validate its robustness, including sensitivity analysis and an ablation study, among other elements.

多标签噪声修正生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。