arXiv:2502.08557cs.IRcs.CL2025-02中稿 · KDD被引 5

用多角色对话式提问提升搜索扩展,避免内容单一。

A New Query Expansion Approach via Agent-Mediated Dialogic Inquiry

  • 三角色协作:提问、答问、反思,层层深化查询
  • 在BEIR和TREC上优于现有方法,检索效果显著提升
  • 适合需要深度理解用户意图的智能搜索场景

查询扩展广泛用于信息检索(IR),通过补充初始查询来改善搜索结果。尽管基于大语言模型(LLM)的方法通过多次提示生成伪相关文本和扩展词,但常导致内容同质化、范围狭窄,缺乏多样上下文以召回相关信息。本文提出AMD:一种代理中介对话框架,包含三个专精角色:(1)苏格拉底式提问代理将初始查询转化为三个子问题,分别基于澄清、假设探查和隐含意义探查等维度;(2)对话式答问代理生成伪答案,从多个视角丰富查询表征,贴合用户意图;(3)反思反馈代理评估并优化这些伪答案,仅保留最相关且有信息量的内容。通过多代理协作机制,AMD有效构建更丰富的查询表征。在BEIR和TREC等多个基准测试中,该框架表现优于现有方法,为检索任务提供了稳健解决方案。

原文摘要 · Abstract (English)

Query expansion is widely used in Information Retrieval (IR) to improve search outcomes by supplementing initial queries with richer information. While recent Large Language Model (LLM) based methods generate pseudo-relevant content and expanded terms via multiple prompts, they often yield homogeneous, narrow expansions that lack the diverse context needed to retrieve relevant information. In this paper, we propose AMD: a new Agent-Mediated Dialogic Framework that engages in a dialogic inquiry involving three specialized roles: (1) a Socratic Questioning Agent reformulates the initial query into three sub-questions, with each question inspired by a specific Socratic questioning dimension, including clarification, assumption probing, and implication probing, (2) a Dialogic Answering Agent generates pseudo-answers, enriching the query representation with multiple perspectives aligned to the user's intent, and (3) a Reflective Feedback Agent evaluates and refines these pseudo-answers, ensuring that only the most relevant and informative content is retained. By leveraging a multi-agent process, AMD effectively crafts richer query representations through inquiry and feedback refinement. Extensive experiments on benchmarks including BEIR and TREC demonstrate that our framework outperforms previous methods, offering a robust solution for retrieval tasks.

信息检索多代理系统查询扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。