通过多样多查询重写提升RAG检索与生成效果
DMQR-RAG: Diverse Multi-Query Rewriting for RAG
- 设计四种不同信息量的查询重写策略,增强文档多样性
- 提出自适应选择机制,减少重写次数并优化整体性能
- 在学术与工业场景均验证有效,提升RAG可靠性
大型语言模型常因静态知识和幻觉问题影响可靠性。检索增强生成(RAG)通过引入外部信息缓解此问题。然而用户查询常含噪声和意图偏差,需查询重写以提高检索文档的相关性。本文提出DMQR-RAG,一种多样多查询重写框架,旨在提升RAG在文档检索与最终响应中的表现。具体而言,研究不同信息量查询对文档多样性的影响,提出四种在不同信息层级运作的重写策略,以增强基线方法性能。此外,设计自适应策略选择方法,在最小化重写次数的同时优化整体表现。所提方法已在学术与工业场景中通过大量实验严格验证。
原文摘要 · Abstract (English)
Large language models often encounter challenges with static knowledge and hallucinations, which undermine their reliability. Retrieval-augmented generation (RAG) mitigates these issues by incorporating external information. However, user queries frequently contain noise and intent deviations, necessitating query rewriting to improve the relevance of retrieved documents. In this paper, we introduce DMQR-RAG, a Diverse Multi-Query Rewriting framework designed to improve the performance of both document retrieval and final responses in RAG. Specifically, we investigate how queries with varying information quantities can retrieve a diverse array of documents, presenting four rewriting strategies that operate at different levels of information to enhance the performance of baseline approaches. Additionally, we propose an adaptive strategy selection method that minimizes the number of rewrites while optimizing overall performance. Our methods have been rigorously validated through extensive experiments conducted in both academic and industry settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。