TREC 2021深度学习赛道更新大规模数据集,验证大模型在检索中的优势。
Overview of the TREC 2021 deep learning track
- 使用更新后的MS MARCO数据集,文档与段落集合规模分别扩大近4倍和16倍。
- 基于预训练的大模型仍显著优于传统检索方法,单阶段模型表现接近多阶段方案。
- 关注新旧数据映射带来的标注完整性与标签质量问题,提出评估反思。
这是TREC深度学习赛道的第三年。延续往年做法,我们采用MS MARCO数据集,为段落与文档排序任务提供数十万条人工标注训练标签。今年对文档与段落语料库进行了全面更新,文档集合规模扩大近4倍,段落集合规模扩大近16倍。基于大规模预训练的深度神经排序模型继续优于传统检索方法。单阶段检索模型在两项任务上均表现良好,但尚未达到多阶段流水线的水平。此外,数据规模增长和整体数据更新引发了对NIST标注完整性的关注,以及旧标签映射至新集合时的标签质量问题,本文对此进行了讨论。
原文摘要 · Abstract (English)
This is the third year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels available for both passage and document ranking tasks. In addition, this year we refreshed both the document and the passage collections which also led to a nearly four times increase in the document collection size and nearly $16$ times increase in the size of the passage collection. Deep neural ranking models that employ large scale pretraininig continued to outperform traditional retrieval methods this year. We also found that single stage retrieval can achieve good performance on both tasks although they still do not perform at par with multistage retrieval pipelines. Finally, the increase in the collection size and the general data refresh raised some questions about completeness of NIST judgments and the quality of the training labels that were mapped to the new collections from the old ones which we discuss in this report.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。