arXiv:2502.20937cs.IR2025-02中稿 · SIGIR 2025被引 11

测试集可能过期:新标注显示部分模型效果下降,提示需更新评测数据。

Variations in Relevance Judgments and the Shelf Life of Test Collections

  • 重标注2019年TREC深度学习赛道数据,检验评估稳定性。
  • 部分神经检索模型在新标注下显著退化,逼近人类排名上限。
  • 提醒长期复用测试集可能导致过拟合,需设定‘过期时间’。

传统Cranfield评估中,系统排序的稳定性不依赖于评估者对单个文档相关性的完全一致。但随着神经检索模型兴起,现代测试集呈现短文档、四等级相关性判断、无描述性信息需求等特点。在此背景下,评估者分歧是否仍可忽略尚不清楚。尤其当这些测试集被反复使用时,更少形式化的信息需求导致更多查询解释可能性,或使测试集产生“过期”。我们开展可复现性研究,重新标注2019年TREC深度学习赛道的相关性判断。结果表明,在神经检索设置下,系统排序仍保持稳定;然而部分模型在新标注下性能显著下降,另一些已达到人类排序水平,暗示测试集存在有效期限。

原文摘要 · Abstract (English)

The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditional test collections. However, the paradigm shift towards neural retrieval models affected the characteristics of modern test collections, e.g., documents are short, judged with four grades of relevance, and information needs have no descriptions or narratives. Under these changes, it is unclear whether assessor disagreement remains negligible for system comparisons. We investigate this aspect under the additional condition that the few modern test collections are heavily re-used. Given more possible query interpretations due to less formalized information needs, an ``expiration date'' for test collections might be needed if top-effectiveness requires overfitting to a single interpretation of relevance. We run a reproducibility study and re-annotate the relevance judgments of the 2019~TREC Deep Learning track. We can reproduce prior work in the neural retrieval setting, showing that assessor disagreement does not affect system rankings. However, we observe that some models substantially degrade with our new relevance judgments, and some have already reached the effectiveness of humans as rankers, providing evidence that test collections can expire.

评测集神经检索过期风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。