arXiv:2505.12694cs.IR2025-05中稿 · SIGIR 2025 short p…被引 17

LLM扩写查询在陌生或模糊问题上会失效,影响检索效果。

LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries

  • 分析了LLM扩写查询在知识不足和歧义下的失败机制。
  • 实验表明知识缺失或歧义高时检索性能显著下降。
  • 为评估此类情况提供新框架,适合检索系统研究者参考。

查询扩写(QE)通过引入相关词提升检索效果,大语言模型(LLMs)提供了传统规则与统计方法的替代方案。然而,基于LLM的QE存在根本缺陷:常无法生成相关知识,导致搜索性能下降。先前研究聚焦于幻觉问题,但其深层原因——LLM知识不足——仍未被充分探讨。本文系统考察两类失败情形:(1) 当LLM缺乏查询知识时,产生错误扩写;(2) 当查询具有歧义时,引发偏倚扩写,缩小搜索覆盖范围。我们在多个数据集上进行受控实验,使用稀疏与密集检索模型评估知识与查询歧义对检索性能的影响。结果表明,当LLM知识不足或查询歧义较高时,基于LLM的QE会显著降低检索有效性。我们提出一种评估框架,用于分析这些条件下的QE表现,为理解基于LLM的检索增强局限性提供洞见。

原文摘要 · Abstract (English)

Query expansion (QE) enhances retrieval by incorporating relevant terms, with large language models (LLMs) offering an effective alternative to traditional rule-based and statistical methods. However, LLM-based QE suffers from a fundamental limitation: it often fails to generate relevant knowledge, degrading search performance. Prior studies have focused on hallucination, yet its underlying cause--LLM knowledge deficiencies--remains underexplored. This paper systematically examines two failure cases in LLM-based QE: (1) when the LLM lacks query knowledge, leading to incorrect expansions, and (2) when the query is ambiguous, causing biased refinements that narrow search coverage. We conduct controlled experiments across multiple datasets, evaluating the effects of knowledge and query ambiguity on retrieval performance using sparse and dense retrieval models. Our results reveal that LLM-based QE can significantly degrade the retrieval effectiveness when knowledge in the LLM is insufficient or query ambiguity is high. We introduce a framework for evaluating QE under these conditions, providing insights into the limitations of LLM-based retrieval augmentation.

检索增强LLM缺陷查询扩写

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。