arXiv:2410.07538cs.LG2024-10

解决众包中完整排序的聚合难题,同时估计问题难度与标注者能力。

Rank Aggregation in Crowdsourcing for Listwise Annotations

  • 基于全局位置信息设计标注质量指标
  • 首次实现无监督同步推断真值排序、标注者能力与问题难度
  • 适用于模型评估和人类反馈强化学习等场景

通过众包进行排名聚合近年来受到广泛关注,尤其在列表式排序标注场景中。然而,现有方法多聚焦单一问题和部分排序,对多个问题中完整列表排序的聚合仍缺乏探索。该场景在模型质量评估和基于人类反馈的强化学习中具有实际意义。为此,我们提出LAC——一种众包环境下的列表式排序聚合方法,通过精确测量并引入全局位置信息。设计了专门的标注质量指标,用于衡量标注排序与真实排序之间的差异,并考虑了排序任务本身的难度,因其直接影响标注者表现并影响最终结果。据我们所知,LAC是首个直接处理列表式众包中完整排序聚合问题的工作,且可无监督地同步推断问题难度、标注者能力与真实排序。为验证方法有效性,我们收集了一个面向实际业务的段落排序数据集。在合成及真实基准数据集上的实验结果表明,所提LAC方法具有显著效果。

原文摘要 · Abstract (English)

Rank aggregation through crowdsourcing has recently gained significant attention, particularly in the context of listwise ranking annotations. However, existing methods primarily focus on a single problem and partial ranks, while the aggregation of listwise full ranks across numerous problems remains largely unexplored. This scenario finds relevance in various applications, such as model quality assessment and reinforcement learning with human feedback. In light of practical needs, we propose LAC, a Listwise rank Aggregation method in Crowdsourcing, where the global position information is carefully measured and included. In our design, an especially proposed annotation quality indicator is employed to measure the discrepancy between the annotated rank and the true rank. We also take the difficulty of the ranking problem itself into consideration, as it directly impacts the performance of annotators and consequently influences the final results. To our knowledge, LAC is the first work to directly deal with the full rank aggregation problem in listwise crowdsourcing, and simultaneously infer the difficulty of problems, the ability of annotators, and the ground-truth ranks in an unsupervised way. To evaluate our method, we collect a real-world business-oriented dataset for paragraph ranking. Experimental results on both synthetic and real-world benchmark datasets demonstrate the effectiveness of our proposed LAC method.

众包排序排名聚合无监督学习标注质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。