LLM让检索基准性能飙升,但可能因数据泄露而失真。
The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and Contamination
- 分析143篇论文,对比TREC Robust04与DL20基准表现变化
- 近年含LLM系统在DL20上nDCG@10提升8.8%,Robust04提升约20%
- 发现两基准存在数据污染,难以判断性能提升是创新还是记忆
基准集合长期推动信息检索(IR)领域的可控比较与累积进步。然而,以往元分析显示,报告的性能提升常无法持续积累,部分原因在于基线过弱或过时。尽管大语言模型(LLM)日益被用于检索流程,其对成熟IR基准的影响尚未系统评估。本研究分析了143篇报告在TREC Robust04和TREC Deep Learning 2020(DL20)段落检索基准上的成果,考察检索有效性与基线强度的长期趋势。我们观察到所谓的「LLM效应」:近期包含LLM组件的系统在DL20上相比TREC 2020最佳结果提升8.8%的nDCG@10,自2023年起在Robust04上提升约20%。然而,采用数据污染检测方法重排序后,发现两个基准均存在可测量的污染。排除污染主题后性能下降,置信区间仍宽,难以判断LLM效应究竟是方法论进步,还是预训练数据中的记忆现象。
原文摘要 · Abstract (English)
Benchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses have shown that reported effectiveness gains often fail to accumulate, in part due to the use of weak or outdated baselines. While large language models are increasingly used in retrieval pipelines, their impact on established IR benchmarks has not been systematically analyzed. In this study, we analyze 143 publications reporting results on the TREC Robust04 collection and the TREC Deep Learning 2020 (DL20) passage retrieval benchmark to examine longitudinal trends in retrieval effectiveness and baseline strength. We observe what we term an \emph{LLM effect}: recent systems incorporating LLM components achieve 8.8\% higher nDCG@10 on DL20 compared to the best result from TREC 2020 and approximately 20\% higher on Robust04 since 2023. However, adapting a data contamination detection approach to reranking reveals measurable contamination in both benchmarks. While excluding contaminated topics reduces effectiveness, confidence intervals remain wide, making it difficult to determine whether the LLM effect reflects genuine methodological advances or memorization from pretraining data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。