arXiv:2504.11284cs.LGcs.AI2025-04ICML被引 2

多标签排序中,标签聚合比损失聚合更不易偏倚。

Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation

  • 用标签聚合替代损失聚合来整合多源标注
  • 损失聚合可能导致某标签主导排序结果
  • 适用于多人标注场景的公平排序建模

二元排序是基础监督学习问题,目标是在单个二值标签下学习使受试者工作特征曲线下面积(AUC)最大化的实例排序。然而,实际中常出现多个二值标签,例如来自不同人工标注者。如何将这些标签融合为一致的排序?本文形式化分析了两种方法——损失聚合与标签聚合,并刻画其贝叶斯最优解。研究表明,尽管两者均可产生帕累托最优解,但损失聚合可能引发标签独裁:无意间过度偏向某一标签。这表明标签聚合优于损失聚合,实验验证了该结论。

原文摘要 · Abstract (English)

Bipartite ranking is a fundamental supervised learning problem, with the goal of learning a ranking over instances with maximal Area Under the ROC Curve (AUC) against a single binary target label. However, one may often observe multiple binary target labels, e.g., from distinct human annotators. How can one synthesize such labels into a single coherent ranking? In this work, we formally analyze two approaches to this problem -- loss aggregation and label aggregation -- by characterizing their Bayes-optimal solutions. We show that while both approaches can yield Pareto-optimal solutions, loss aggregation can exhibit label dictatorship: one can inadvertently (and undesirably) favor one label over others. This suggests that label aggregation can be preferable to loss aggregation, which we empirically verify.

排序学习多标签模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。