ConRAG通过多视角证据融合提升复杂多跳问答性能。
ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering

- 从查询与语料双端优化,融合关系、实体、文本多视图信号
- 在三个基准上超越基线,最高提升26.9%,刷新MuSiQue记录
- 适合需要精准多步推理的问答系统研发者
检索增强生成(RAG)已成为提升大语言模型在多跳问答任务中表现的有前景范式,该任务需基于多文档证据进行推理。现有方法通常聚焦于查询端的任务分解或语料端的知识图谱构建,但仍在复杂多跳问答任务中表现不足。为此,我们提出ConRAG——一种共识驱动的多视角RAG框架,通过系统优化查询与语料两端,并利用关系、实体和文本多视图证据实现更精准检索,显著提升大模型性能。在三个多跳问答基准上的大量实验表明,ConRAG始终以明显优势超越所有基线,例如相较于原始RAG平均性能提升高达26.9%,并使Gemma-4-31B在具有挑战性的MuSiQue基准上达到新的最佳水平。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA), which requires reasoning over evidence from multiple documents. Current multi-hop RAG methods generally focus on either query-side task decomposition or corpus-side knowledge graph construction. Despite their progress, these methods still struggle to achieve satisfactory performance on complex multi-hop QA tasks. To this end, we propose ConRAG, a consensus-driven multi-view RAG framework that effectively boosts LLMs on complex multi-hop QA. The core of ConRAG is to systematically optimize both the query and corpus sides and to leverage multi-view evidence (relation, entity, and text signals) for more accurate retrieval. Extensive experiments on three multi-hop QA benchmarks show that ConRAG consistently outperforms all baselines by a clear margin, e.g., up to +26.9% average performance gains over vanilla RAG, and enables Gemma-4-31B to achieve a new state-of-the-art record on the challenging MuSiQue benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。