让多标签排序学会区分正类重要性,提升排序精度。
UniMLR: Modeling Implicit Class Significance for Multi-Label Ranking
- 用正类内部排序建模隐含重要性分布,替代传统等权重假设。
- 在真实与合成数据上验证,模型能准确学习正类排序顺序。
- 适用于标注稀疏或存在偏见的多标签排序任务。
现有多标签排序框架仅依赖标签的正负划分,未能利用正类间的排序信息。本文提出UniMLR,一种新范式,通过正类间的排序关系建模隐含类别相关性/重要性值的概率分布,而非将其视为同等重要。该方法统一了多标签排序与分类任务。为应对多标签排序数据集中的标注稀缺与偏差问题,我们引入八个合成数据集(Ranked MNISTs),其生成受不同显著性决定因素影响,构建出丰富且可控的实验环境。统计结果表明,本方法能准确学习正类排序顺序,且与真实情况一致,并与潜在显著性值成比例。我们在真实与合成数据集上进行全面实验,验证了所提框架的有效性。
原文摘要 · Abstract (English)
Existing multi-label ranking (MLR) frameworks only exploit information deduced from the bipartition of labels into positive and negative sets. Therefore, they do not benefit from ranking among positive labels, which is the novel MLR approach we introduce in this paper. We propose UniMLR, a new MLR paradigm that models implicit class relevance/significance values as probability distributions using the ranking among positive labels, rather than treating them as equally important. This approach unifies ranking and classification tasks associated with MLR. Additionally, we address the challenges of scarcity and annotation bias in MLR datasets by introducing eight synthetic datasets (Ranked MNISTs) generated with varying significance-determining factors, providing an enriched and controllable experimental environment. We statistically demonstrate that our method accurately learns a representation of the positive rank order, which is consistent with the ground truth and proportional to the underlying significance values. Finally, we conduct comprehensive empirical experiments on both real-world and synthetic datasets, demonstrating the value of our proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。