TREC 2022深挖海量标注数据,强化了文档与段落检索评测体系。
Overview of the TREC 2022 deep learning track
- 基于更新的MS MARCO数据集,段落与文档集合规模分别扩大16倍和4倍
- 大模型预训练方法持续领先传统检索,但单阶段稠密检索表现不如去年
- 聚焦段落检索测试集构建,提升评测质量,支持未来复用
TREC 2022是深度学习赛道的第四年。沿用往年使用的MS MARCO数据集,提供数十万条人工标注的段落与文档排序标签。今年还引入去年发布的更新版段落与文档集合,使段落集合规模扩大近16倍,文档集合规模扩大近4倍。与往年不同,2022年主要致力于构建更完整的段落检索测试集,文档排序任务作为次要任务,其文档级标签由段落级标签推断得出。分析显示,类似往年,采用大规模预训练的深度神经排序模型仍优于传统方法。由于将评估资源集中于段落标注,本年度查询与判断质量更可靠,且能更好区分各系统表现,并支持未来重复使用该数据集。部分结果出人意料:一些顶尖表现系统未采用稠密检索,而单阶段稠密检索系统的竞争力较去年下降。
原文摘要 · Abstract (English)
This is the fourth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels available for both passage and document ranking tasks. In addition, this year we also leverage both the refreshed passage and document collections that were released last year leading to a nearly $16$ times increase in the size of the passage collection and nearly four times increase in the document collection size. Unlike previous years, in 2022 we mainly focused on constructing a more complete test collection for the passage retrieval task, which has been the primary focus of the track. The document ranking task was kept as a secondary task, where document-level labels were inferred from the passage-level labels. Our analysis shows that similar to previous years, deep neural ranking models that employ large scale pretraining continued to outperform traditional retrieval methods. Due to the focusing our judging resources on passage judging, we are more confident in the quality of this year's queries and judgments, with respect to our ability to distinguish between runs and reuse the dataset in future. We also see some surprises in overall outcomes. Some top-performing runs did not do dense retrieval. Runs that did single-stage dense retrieval were not as competitive this year as they were last year.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。