arXiv:2505.17875cs.LG2025-05被引 1

解决多标签半监督特征选择中的标签相关性与图结构问题

Semi-Supervised Multi-Label Feature Selection with Consistent Sparse Graph Learning

  • 通过低维标签子空间学习标签相关性,避免特征被误删
  • 联合重构标签空间与子空间,自适应构建一致性相似图
  • 适合标签复杂、标注样本少的高维数据特征筛选场景

在实际应用中,高维数据常伴随多种语义标签,但传统特征选择方法多针对单标签设计。现有半监督多标签方法面临两大挑战:(1) 标注样本不足时难以评估标签相关性,导致特定标签特征被错误剔除;(2) 现有图方法基于原始特征空间构建的相似性图不适用于多标签问题,引发不可靠软标签及性能下降。为此,提出一致稀疏图学习的多标签半监督特征选择方法(SGMFS),通过从投影特征中学习低维独立标签子空间,兼容多标签并有效捕捉标签相关性;同时,在标签空间与学习子空间中同步进行稀疏重构,自适应学习保持一致性结构的相似图,促进未标记样本的合理软标签传播,提升特征选择性能。设计了快速收敛的优化算法,大量实验验证了其优越性。

原文摘要 · Abstract (English)

In practical domains, high-dimensional data are usually associated with diverse semantic labels, whereas traditional feature selection methods are designed for single-label data. Moreover, existing multi-label methods encounter two main challenges in semi-supervised scenarios: (1). Most semi-supervised methods fail to evaluate the label correlations without enough labeled samples, which are the critical information of multi-label feature selection, making label-specific features discarded. (2). The similarity graph structure directly derived from the original feature space is suboptimal for multi-label problems in existing graph-based methods, leading to unreliable soft labels and degraded feature selection performance. To overcome them, we propose a consistent sparse graph learning method for multi-label semi-supervised feature selection (SGMFS), which can enhance the feature selection performance by maintaining space consistency and learning label correlations in semi-supervised scenarios. Specifically, for Challenge (1), SGMFS learns a low-dimensional and independent label subspace from the projected features, which can compatibly cross multiple labels and effectively achieve the label correlations. For Challenge (2), instead of constructing a fixed similarity graph for semi-supervised learning, SGMFS thoroughly explores the intrinsic structure of the data by performing sparse reconstruction of samples in both the label space and the learned subspace simultaneously. In this way, the similarity graph can be adaptively learned to maintain the consistency between label space and the learned subspace, which can promote propagating proper soft labels for unlabeled samples, facilitating the ultimate feature selection. An effective solution with fast convergence is designed to optimize the objective function. Extensive experiments validate the superiority of SGMFS.

多标签半监督特征选择图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。