arXiv:2601.11722cs.CLcs.IR2026-01被引 1

让对话搜索的追问问题有据可依,避免无中生有。

RAC: Retrieval-Augmented Clarification for Faithful Conversational Search

  • 用检索增强生成基于文档的追问问题
  • 在四个基准上显著提升追问忠实度
  • 适合需要可靠对话问答的场景

澄清问题能帮助对话式搜索系统解决用户查询中的模糊或不完整问题。以往工作主要关注问题的流畅性和与用户意图的一致性,尤其通过特征提取实现,但较少关注澄清问题是否基于底层语料库。缺乏这种依据会导致系统提出无法从现有文档中回答的问题。本文提出RAC(检索增强澄清)框架,生成基于语料库的忠实澄清问题。在比较多种索引策略后,微调大语言模型以充分利用研究上下文,并鼓励生成有证据支持的问题。随后采用对比偏好优化,使被检索段落支持的问题优于无根据的替代方案。在四个基准上的评估表明,RAC相比基线有显著提升。除了使用大模型作为裁判的评估外,还引入基于NLI和数据到文本的新指标,评估问题与上下文的锚定程度,结果证明该方法持续提升了忠实度。

原文摘要 · Abstract (English)

Clarification questions help conversational search systems resolve ambiguous or underspecified user queries. While prior work has focused on fluency and alignment with user intent, especially through facet extraction, much less attention has been paid to grounding clarifications in the underlying corpus. Without such grounding, systems risk asking questions that cannot be answered from the available documents. We introduce RAC (Retrieval-Augmented Clarification), a framework for generating corpus-faithful clarification questions. After comparing several indexing strategies for retrieval, we fine-tune a large language model to make optimal use of research context and to encourage the generation of evidence-based question. We then apply contrastive preference optimization to favor questions supported by retrieved passages over ungrounded alternatives. Evaluated on four benchmarks, RAC demonstrate significant improvements over baselines. In addition to LLM-as-Judge assessments, we introduce novel metrics derived from NLI and data-to-text to assess how well questions are anchored in the context, and we demonstrate that our approach consistently enhances faithfulness.

对话搜索检索增强忠实生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。