arXiv:2603.11008cs.IRcs.CL2026-03被引 1

探究大模型伪相关反馈中两个关键设计的选择影响。

A Systematic Study of Pseudo-Relevance Feedback with LLMs

  • 分离分析反馈来源与反馈模型的作用,控制实验验证
  • 仅用大模型生成文本作反馈成本最低且效果好
  • 强检索器配合语料反馈时效果最佳,适合低资源场景

基于大语言模型(LLMs)的伪相关反馈(PRF)方法可从两个关键设计维度进行组织:反馈来源(反馈文本来自何处)和反馈模型(如何利用反馈文本优化查询表示)。然而,这两个维度在现有评估中常被混淆,其独立作用尚不明确。本文通过受控实验,系统研究了反馈来源与反馈模型对PRF效果的影响。在13个低资源BEIR任务上,对比五种LLM PRF方法,结果表明:(1) 反馈模型的选择对PRF效果具有决定性影响;(2) 仅使用大模型生成的文本作为反馈是最具成本效益的方案;(3) 当结合强第一阶段检索器的候选文档时,来自语料的反馈最具优势。研究揭示了PRF设计空间中真正关键的要素。

原文摘要 · Abstract (English)

Pseudo-relevance feedback (PRF) methods built on large language models (LLMs) can be organized along two key design dimensions: the feedback source, which is where the feedback text is derived from and the feedback model, which is how the given feedback text is used to refine the query representation. However, the independent role that each dimension plays is unclear, as both are often entangled in empirical evaluations. In this paper, we address this gap by systematically studying how the choice of feedback source and feedback model impact PRF effectiveness through controlled experimentation. Across 13 low-resource BEIR tasks with five LLM PRF methods, our results show: (1) the choice of feedback model can play a critical role in PRF effectiveness; (2) feedback derived solely from LLM-generated text provides the most cost-effective solution; and (3) feedback derived from the corpus is most beneficial when utilizing candidate documents from a strong first-stage retriever. Together, our findings provide a better understanding of which elements in the PRF design space are most important.

信息检索大模型伪相关反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。