arXiv:2503.00674cs.IRcs.AI2025-03被引 4

构建首个用于自然语言处理的有序相关性评估基准,提升排序模型精细区分能力。

OrdRankBen: A Novel Ranking Benchmark for Ordinal Relevance in NLP

  • 引入结构化有序标签,突破传统二值或连续评分局限。
  • 在两个不同有序分布的数据集上验证,显著提升模型区分细粒度相关性的能力。
  • 适合需要精准排序的场景,如信息检索、推荐系统等研究者使用。

自然语言处理中的排序任务评估仍面临重大挑战,尤其因真实场景中缺乏直接的结果标签。基准数据集在提供标准化测试平台、确保公平比较、提高可复现性及追踪进展方面至关重要。现有NLP排序基准多采用二值相关性标签或连续相关性得分,忽视了有序相关性。二值标签过于简化相关性差异,而连续得分缺乏明确的有序结构,难以有效捕捉细微的排序区别。为此,我们提出OrdRankBen,一个旨在捕捉多粒度相关性差异的新基准。不同于传统基准,OrdRankBen采用结构化有序标签,实现更精确的排序评估。由于缺乏适用于有序相关性排序的NLP数据集,我们构建了两个具有不同有序标签分布的数据集。进一步对三类模型(基于排序的语言模型、通用大语言模型、专注排序的大语言模型)进行了评估。实验结果表明,有序相关性建模能更精确地评估排序模型,增强其对排序项间多粒度差异的分辨能力,这对需要细粒度相关性区分的任务至关重要。

原文摘要 · Abstract (English)

The evaluation of ranking tasks remains a significant challenge in natural language processing (NLP), particularly due to the lack of direct labels for results in real-world scenarios. Benchmark datasets play a crucial role in providing standardized testbeds that ensure fair comparisons, enhance reproducibility, and enable progress tracking, facilitating rigorous assessment and continuous improvement of ranking models. Existing NLP ranking benchmarks typically use binary relevance labels or continuous relevance scores, neglecting ordinal relevance scores. However, binary labels oversimplify relevance distinctions, while continuous scores lack a clear ordinal structure, making it challenging to capture nuanced ranking differences effectively. To address these challenges, we introduce OrdRankBen, a novel benchmark designed to capture multi-granularity relevance distinctions. Unlike conventional benchmarks, OrdRankBen incorporates structured ordinal labels, enabling more precise ranking evaluations. Given the absence of suitable datasets for ordinal relevance ranking in NLP, we constructed two datasets with distinct ordinal label distributions. We further evaluate various models for three model types, ranking-based language models, general large language models, and ranking-focused large language models on these datasets. Experimental results show that ordinal relevance modeling provides a more precise evaluation of ranking models, improving their ability to distinguish multi-granularity differences among ranked items-crucial for tasks that demand fine-grained relevance differentiation.

排序评估有序标签基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。