arXiv:2410.13579cs.LG2024-10被引 1

解决标签分布不均衡问题,提升不完整标签学习性能

Towards Better Performance in Incomplete LDL: Addressing Data Imbalance

  • 将标签分布分解为低秩(高频)与稀疏(低频)两部分建模
  • 在15个真实数据集上显著优于现有不完整标签方法
  • 理论证明泛化误差有界,适合标签分布不均场景

标签分布学习(LDL)是一种应对标签模糊性的新范式,在实际应用中广泛存在。然而,真实场景中获取完整的标签分布极具挑战,催生了不完整标签分布学习(InLDL)。现有InLDL方法忽略了标签分布的内在不平衡性。为此,本文提出一种新框架——不完整且不平衡标签分布学习(I²LDL),同时处理不完整标签与不平衡标签分布。该方法将标签分布矩阵分解为低秩成分(用于常见标签)和稀疏成分(用于罕见标签),有效捕捉头尾标签结构。通过交替方向乘子法(ADMM)优化模型,并利用Rademacher复杂度推导泛化误差界,提供强理论保证。在15个真实世界数据集上的大量实验表明,所提框架在性能与鲁棒性上均优于现有InLDL方法。

原文摘要 · Abstract (English)

Label Distribution Learning (LDL) is a novel machine learning paradigm that addresses the problem of label ambiguity and has found widespread applications. Obtaining complete label distributions in real-world scenarios is challenging, which has led to the emergence of Incomplete Label Distribution Learning (InLDL). However, the existing InLDL methods overlook a crucial aspect of LDL data: the inherent imbalance in label distributions. To address this limitation, we propose \textbf{Incomplete and Imbalance Label Distribution Learning (I\(^2\)LDL)}, a framework that simultaneously handles incomplete labels and imbalanced label distributions. Our method decomposes the label distribution matrix into a low-rank component for frequent labels and a sparse component for rare labels, effectively capturing the structure of both head and tail labels. We optimize the model using the Alternating Direction Method of Multipliers (ADMM) and derive generalization error bounds via Rademacher complexity, providing strong theoretical guarantees. Extensive experiments on 15 real-world datasets demonstrate the effectiveness and robustness of our proposed framework compared to existing InLDL methods.

标签分布学习不均衡数据机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。