用思维演化机制提升搜索关键词扩展的全面性与多样性
ThinkQE: Query Expansion via an Evolving Thinking Process
- 引入思维演化过程,促进深层语义探索
- 通过检索反馈迭代优化扩展词,提升结果多样性
- 无需训练即可超越主流稠密检索模型,适合实际部署
有效的网络搜索查询扩展需要兼顾探索深度与结果多样性,以捕捉查询的多重含义和方面。尽管近期基于大语言模型的方法在不需额外训练的情况下提升了检索性能并具备强领域泛化能力,但常生成过于聚焦的扩展词,忽视上述目标。本文提出ThinkQE,一种测试时查询扩展框架,包含两个核心组件:基于思维的扩展过程,推动更深入全面的语义探索;以及基于语料库交互的策略,利用来自语料库的检索反馈迭代优化扩展内容。在多个网络搜索基准(DL19、DL20 和 BRIGHT)上的实验表明,ThinkQE 持续优于先前方法,包括依赖训练的稠密检索器和重排序器。
原文摘要 · Abstract (English)
Effective query expansion for web search benefits from promoting both exploration and result diversity to capture multiple interpretations and facets of a query. While recent LLM-based methods have improved retrieval performance and demonstrate strong domain generalization without additional training, they often generate narrowly focused expansions that overlook these desiderata. We propose ThinkQE, a test-time query expansion framework addressing this limitation through two key components: a thinking-based expansion process that encourages deeper and comprehensive semantic exploration, and a corpus-interaction strategy that iteratively refines expansions using retrieval feedback from the corpus. Experiments on diverse web search benchmarks (DL19, DL20, and BRIGHT) show ThinkQE consistently outperforms prior approaches, including training-intensive dense retrievers and rerankers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。