用检索增强生成澄清问题,让模型提问更贴合实际信息。
Corpus-informed Retrieval Augmented Generation of Clarifying Questions
- 用RAG联合建模用户查询与检索文档,自动发现疑问点并生成澄清问题。
- 提升证据文档数量可拓宽问题覆盖范围,但现有数据多不支持真实意图。
- 提出数据增强方法对齐问题与文档内容,适合搜索系统优化研究者。
本研究旨在开发能够生成基于语料库的澄清问题的模型,确保问题与检索语料库中的信息一致。实验表明,检索增强语言模型(RAG)在该任务中表现有效,其优势在于:(i) 能联合建模用户查询与检索语料库,端到端地定位不确定性并提出澄清问题;(ii) 可建模更多证据文档,从而扩展问题的广度。然而,我们发现当前数据集中多数搜索意图并未被语料库支持,这对训练和评估造成困扰。这导致问题生成模型产生幻觉——提出语料库中不存在的意图,严重影响性能。为此,我们提出数据增强方法,使真实澄清问题与检索语料库对齐。此外,探索了推理阶段提升证据相关性的技术,但仍难以在语料库中准确识别真实意图。分析表明,这一挑战部分源于现有数据集对澄清分类体系的偏向,亟需支持语料库导向澄清生成的数据集。
原文摘要 · Abstract (English)
This study aims to develop models that generate corpus informed clarifying questions for web search, in a way that ensures the questions align with the available information in the retrieval corpus. We demonstrate the effectiveness of Retrieval Augmented Language Models (RAG) in this process, emphasising their ability to (i) jointly model the user query and retrieval corpus to pinpoint the uncertainty and ask for clarifications end-to-end and (ii) model more evidence documents, which can be used towards increasing the breadth of the questions asked. However, we observe that in current datasets search intents are largely unsupported by the corpus, which is problematic both for training and evaluation. This causes question generation models to ``hallucinate'', ie. suggest intents that are not in the corpus, which can have detrimental effects in performance. To address this, we propose dataset augmentation methods that align the ground truth clarifications with the retrieval corpus. Additionally, we explore techniques to enhance the relevance of the evidence pool during inference, but find that identifying ground truth intents within the corpus remains challenging. Our analysis suggests that this challenge is partly due to the bias of current datasets towards clarification taxonomies and calls for data that can support generating corpus-informed clarifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。