提出噪声鲁棒的特征选择方法,提升部分多标签学习中正标签识别能力。
Noise-Resistant Label Reconstruction Feature Selection for Partial Multi-Label Learning
- 利用标签关系的抗噪性消除标签歧义
- 基于改进低秩假设重构权重矩阵,保留高维信息
- 适合标签噪声大、正标签难辨的现实场景
维度灾难在各类数据模式中普遍存在,易导致模型过拟合并降低分类性能。然而,现有针对部分多标签学习(PML)的研究较少关注此问题。在PML中,每个样本关联一组候选标签,其中至少一个为正确标签。现有方法多基于低秩假设,但该假设在实际中难以满足,且可能丢失高维信息。此外,我们发现现有方法对正标签的识别能力较差,这在真实场景中尤为关键。本文提出一种考虑数据集两个重要特性的PML特征选择方法:标签关系的抗噪性与标签连通性。利用标签关系的抗噪性进行标签消歧,通过改进的低秩假设设计学习过程。最终通过标签连通性找出代表性标签,并重构权重矩阵,筛选出对这些标签具有强识别能力的特征。在基准数据集上的实验结果表明,所提方法具有显著优势。
原文摘要 · Abstract (English)
The "Curse of dimensionality" is prevalent across various data patterns, which increases the risk of model overfitting and leads to a decline in model classification performance. However, few studies have focused on this issue in Partial Multi-label Learning (PML), where each sample is associated with a set of candidate labels, at least one of which is correct. Existing PML methods addressing this problem are mainly based on the low-rank assumption. However, low-rank assumption is difficult to be satisfied in practical situations and may lead to loss of high-dimensional information. Furthermore, we find that existing methods have poor ability to identify positive labels, which is important in real-world scenarios. In this paper, a PML feature selection method is proposed considering two important characteristics of dataset: label relationship's noise-resistance and label connectivity. Our proposed method utilizes label relationship's noise-resistance to disambiguate labels. Then the learning process is designed through the reformed low-rank assumption. Finally, representative labels are found through label connectivity, and the weight matrix is reconstructed to select features with strong identification ability to these labels. The experimental results on benchmark datasets demonstrate the superiority of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。