通过硬标签重构提升有偏标注下的标签分布学习效果
Label Distribution Learning with Biased Annotations by Learning Multi-Label Representation
- 将软标签转为硬多标签,降低噪声干扰
- 在真实数据集上准确恢复真实标签分布
- 适合处理有偏差标注的多标签场景
多标签学习(MLL)因其能有效表示现实世界数据而受到关注。标签分布学习(LDL)作为其扩展,致力于从标签分布中学习,但准确获取标签分布仍具挑战。现有方法基于低秩假设,通过探索标签相关性从有偏观测中恢复真实分布。然而,最新研究表明标签分布往往为满秩,对有偏观测直接应用低秩近似会导致恢复不准确并引发性能下降。本文提出新思路:先将软标签分布退化为硬多热编码标签,再恢复每个样本的真实标签信息。该方法源于一个观察——分配硬标签比软标签更易且更具抗噪性,从而减小标签偏差。同时,假设预测标签分布的多标签空间具有低秩特性,可更合理地捕捉标签相关性。理论分析与实验验证了该方法在真实数据集上的有效性与鲁棒性。
原文摘要 · Abstract (English)
Multi-label learning (MLL) has gained attention for its ability to represent real-world data. Label Distribution Learning (LDL), an extension of MLL to learning from label distributions, faces challenges in collecting accurate label distributions. To address the issue of biased annotations, based on the low-rank assumption, existing works recover true distributions from biased observations by exploring the label correlations. However, recent evidence shows that the label distribution tends to be full-rank, and naive apply of low-rank approximation on biased observation leads to inaccurate recovery and performance degradation. In this paper, we address the LDL with biased annotations problem from a novel perspective, where we first degenerate the soft label distribution into a hard multi-hot label and then recover the true label information for each instance. This idea stems from an insight that assigning hard multi-hot labels is often easier than assigning a soft label distribution, and it shows stronger immunity to noise disturbances, leading to smaller label bias. Moreover, assuming that the multi-label space for predicting label distributions is low-rank offers a more reasonable approach to capturing label correlations. Theoretical analysis and experiments confirm the effectiveness and robustness of our method on real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。