arXiv:2503.13874cs.LG2025-03被引 2

用二值哈希编码提升多标签特征选择的可靠性,减少噪声干扰。

Multi-label feature selection based on binary hashing learning and dynamic graph constraints

  • 用低维二值哈希码替代连续伪标签,降低噪声。
  • 构建动态图约束空间,增强样本间关系建模的可靠性。
  • 适合需要高鲁棒性特征选择的多标签场景,如文本分类。

多标签学习在从标签空间提取可靠监督信号方面面临挑战。现有方法常使用连续伪标签替代二值标签以增强监督信息表示,但可能引入无关标签的噪声并导致不可靠的图结构。为此,本文提出首个将二值哈希融入多标签学习的方法——二值哈希与动态图约束(BHDG)。BHDG利用低维二值哈希码作为伪标签,减少噪声并提升表示鲁棒性;基于这些伪标签的图结构,构建动态约束的样本投影空间,增强动态图可靠性。为进一步优化伪标签质量,BHDG在样本空间中引入标签图约束和内积最小化,并在目标函数中加入$l_{2,1}$-范数正则项以促进特征选择。采用增广拉格朗日乘子法有效优化二值变量。在10个基准数据集上的全面实验表明,BHDG在六个评估指标上均优于十种前沿方法,整体性能排名第一,平均超越次优方法至少2.7个排名,验证了其在多标签特征选择中的有效性与鲁棒性。

原文摘要 · Abstract (English)

Multi-label learning poses significant challenges in extracting reliable supervisory signals from the label space. Existing approaches often employ continuous pseudo-labels to replace binary labels, improving supervisory information representation. However, these methods can introduce noise from irrelevant labels and lead to unreliable graph structures. To overcome these limitations, this study introduces a novel multi-label feature selection method called Binary Hashing and Dynamic Graph Constraint (BHDG), the first method to integrate binary hashing into multi-label learning. BHDG utilizes low-dimensional binary hashing codes as pseudo-labels to reduce noise and improve representation robustness. A dynamically constrained sample projection space is constructed based on the graph structure of these binary pseudo-labels, enhancing the reliability of the dynamic graph. To further enhance pseudo-label quality, BHDG incorporates label graph constraints and inner product minimization within the sample space. Additionally, an $l_{2,1}$-norm regularization term is added to the objective function to facilitate the feature selection process. The augmented Lagrangian multiplier (ALM) method is employed to optimize binary variables effectively. Comprehensive experiments on 10 benchmark datasets demonstrate that BHDG outperforms ten state-of-the-art methods across six evaluation metrics. BHDG achieves the highest overall performance ranking, surpassing the next-best method by an average of at least 2.7 ranks per metric, underscoring its effectiveness and robustness in multi-label feature selection.

多标签学习特征选择二值哈希图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。