修正标签缺失设定,让模型从不完整标签中恢复隐藏标签
Towards Better IncomLDL: We Are Unaware of Hidden Labels in Advance
- 用比例约束捕捉可见标签的分布信息
- 结合局部相似性和全局低秩结构推理隐藏标签
- 理论证明可恢复性,适合真实场景的不完整标注
标签分布学习(LDL)通过样本的标签分布描述其语义。然而获取完整标签分布数据集成本高昂,催生了不完整标签分布学习(IncomLDL)。现有方法将缺失标签的置信度设为0,其余标签不变,但此设定不现实——当某些标签缺失时,剩余标签的置信度应相应提升。本文纠正该假设,提出新问题:隐藏标签下的标签分布学习(HidLDL),旨在从真实世界中被遗漏部分标签的不完整分布中恢复完整分布。为此,我们发现可见标签的比例信息至关重要,引入创新约束在优化中利用该信息;同时结合局部特征相似性与全局低秩结构,揭示隐藏标签的潜在模式。此外,我们从理论上给出了方法的恢复边界,证明其可行性。在多个数据集上的恢复与预测实验表明,本方法显著优于当前最先进的LDL与IncomLDL方法。
原文摘要 · Abstract (English)
Label distribution learning (LDL) is a novel paradigm that describe the samples by label distribution of a sample. However, acquiring LDL dataset is costly and time-consuming, which leads to the birth of incomplete label distribution learning (IncomLDL). All the previous IncomLDL methods set the description degrees of "missing" labels in an instance to 0, but remains those of other labels unchanged. This setting is unrealistic because when certain labels are missing, the degrees of the remaining labels will increase accordingly. We fix this unrealistic setting in IncomLDL and raise a new problem: LDL with hidden labels (HidLDL), which aims to recover a complete label distribution from a real-world incomplete label distribution where certain labels in an instance are omitted during annotation. To solve this challenging problem, we discover the significance of proportional information of the observed labels and capture it by an innovative constraint to utilize it during the optimization process. We simultaneously use local feature similarity and the global low-rank structure to reveal the mysterious veil of hidden labels. Moreover, we theoretically give the recovery bound of our method, proving the feasibility of our method in learning from hidden labels. Extensive recovery and predictive experiments on various datasets prove the superiority of our method to state-of-the-art LDL and IncomLDL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。