大规模文档重排时,现有重排模型效果会下降甚至变差。
Drowning in Documents: Consequences of Scaling Reranker Inference
- 用强初检系统测试重排器在多文档场景下的表现
- 重排效果随文档数增加先升后降,超限时反而更差
- 适合关注检索效率与质量平衡的研究者
重排器(通常为交叉编码器)计算成本高,但普遍被认为优于廉价的初始信息检索系统。我们通过在完整检索任务上评估重排器性能(而非仅对一阶段检索结果重新打分),挑战这一假设。为获得更可靠的评估,我们采用现代密集嵌入构建强初检系统,并在多种精心设计的挑战性任务上测试重排器,包括内部构建数据集以避免污染以及跨领域数据集。实证结果揭示了一个意外趋势:最佳现有重排器在逐步增加待评分文档数量时初期有提升,但其有效性逐渐下降,甚至在超过某一阈值后导致检索质量退化。我们希望这些发现能推动未来重排技术的改进。
原文摘要 · Abstract (English)
Rerankers, typically cross-encoders, are computationally intensive but are frequently used because they are widely assumed to outperform cheaper initial IR systems. We challenge this assumption by measuring reranker performance for full retrieval, not just re-scoring first-stage retrieval. To provide a more robust evaluation, we prioritize strong first-stage retrieval using modern dense embeddings and test rerankers on a variety of carefully chosen, challenging tasks, including internally curated datasets to avoid contamination, and out-of-domain ones. Our empirical results reveal a surprising trend: the best existing rerankers provide initial improvements when scoring progressively more documents, but their effectiveness gradually declines and can even degrade quality beyond a certain limit. We hope that our findings will spur future research to improve reranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。