arXiv:2509.11353cs.IR2025-09被引 18

大模型在排序中偏好新文档,可能误导信息检索结果。

Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking

  • 通过伪造发布时间测试大模型排序倾向
  • 新文档平均排名提升95位,年份前移4.78年
  • 即使大模型也难避免,适合关注公平性的研究者

大语言模型(LLMs)被广泛用于信息检索系统的第二阶段重排,但其对时效性的偏好尚未受到足够关注。本研究在TREC Deep Learning 2021(DL21)和2022(DL22)语料库中,人为添加出版日期,测试了七种模型(GPT-3.5-turbo、GPT-4o、GPT-4、LLaMA-3 8B/70B、Qwen-2.5 7B/72B)对较新文档的隐式偏好。在列表级重排实验中,'新鲜'文档的Top-10平均发布时间提前最多4.78年,个别条目排名上升高达95位。在成对偏好实验中,当两篇相关性相同的文档仅因时间不同而被注入日期后,模型偏好反转平均达25%。尽管模型规模越大,偏倚越小,但无一能完全消除该现象。结果表明,大模型普遍存在显著的时效性偏差,亟需有效的缓解策略。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in information systems, including being used as second-stage rerankers in information retrieval pipelines, yet their susceptibility to recency bias has received little attention. We investigate whether LLMs implicitly favour newer documents by prepending artificial publication dates to passages in the TREC Deep Learning passage retrieval collections in 2021 (DL21) and 2022 (DL22). Across seven models, GPT-3.5-turbo, GPT-4o, GPT-4, LLaMA-3 8B/70B, and Qwen-2.5 7B/72B, "fresh" passages are consistently promoted, shifting the Top-10's mean publication year forward by up to 4.78 years and moving individual items by as many as 95 ranks in our listwise reranking experiments. Although larger models attenuate the effect, none eliminate it. We also observe that the preference of LLMs between two passages with an identical relevance level can be reversed by up to 25% on average after date injection in our pairwise preference experiments. These findings provide quantitative evidence of a pervasive recency bias in LLMs and highlight the importance of effective bias-mitigation strategies.

大模型排序偏差信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。