用大模型自动生成文档提升问答效果,发现有效类型和组合方式。
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models
- 基于系统功能语言学构建自生成文档分类体系
- 识别出对问答任务最有帮助的文档类型
- 提供可落地的文档融合策略,适合知识密集型任务
将大语言模型自动生成的文档(Self-Docs)与外部检索文档结合,已成为增强检索增强生成(RAG)系统的有效方法。然而,以往研究多聚焦于如何优化使用自生成文档,对其内在特性仍缺乏深入探索。为弥补这一空白,我们首先评估自生成文档的整体有效性,识别影响其在RAG中表现的关键因素(RQ1)。在此基础上,我们基于系统功能语言学构建分类体系,对比不同类别自生成文档的影响(RQ2),并探究其与外部来源的融合策略(RQ3)。研究发现,特定类型的自生成文档最有助于提升性能,并提出了实用指南,可在知识密集型问答任务中实现显著改进。
原文摘要 · Abstract (English)
The integration of documents generated by LLMs themselves (Self-Docs) alongside retrieved documents has emerged as a promising strategy for retrieval-augmented generation systems. However, previous research primarily focuses on optimizing the use of Self-Docs, with their inherent properties remaining underexplored. To bridge this gap, we first investigate the overall effectiveness of Self-Docs, identifying key factors that shape their contribution to RAG performance (RQ1). Building on these insights, we develop a taxonomy grounded in Systemic Functional Linguistics to compare the influence of various Self-Docs categories (RQ2) and explore strategies for combining them with external sources (RQ3). Our findings reveal which types of Self-Docs are most beneficial and offer practical guidelines for leveraging them to achieve significant improvements in knowledge-intensive question answering tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。