研究检索质量能否预示生成内容的信息覆盖度,发现二者有强相关性。
Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage
- 通过多组检索-生成实验,验证检索指标与生成信息覆盖率的关联性。
- 在多个基准上,检索效果与生成结果的信息覆盖度相关性达0.7以上。
- 适合关注RAG系统优化和评估方法的研究者参考。
检索增强生成(RAG)系统结合文档检索与生成模型,用于报告生成等复杂信息任务。尽管检索质量与生成效果的关系看似直观,但尚未系统研究。本文在两个文本RAG基准(TREC NeuCLIR 2024、TREC RAG 2024)和一个跨模态基准(WikiVideo)上,分析了15个文本检索堆栈和10个跨模态检索堆栈,涵盖四种RAG流程和多种评估框架(Auto-ARGUE与MiRAGE)。结果表明,在主题级和系统级上,基于覆盖率的检索指标与生成响应中的信息块覆盖率存在强相关性。当检索目标与生成目标对齐时,相关性最强;而更复杂的迭代式RAG流程会部分削弱这一关联。研究为使用检索指标作为RAG性能代理提供了实证支持。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) systems combine document retrieval with a generative model to address complex information seeking tasks like report generation. While the relationship between retrieval quality and generation effectiveness seems intuitive, it has not been systematically studied. We investigate whether upstream retrieval metrics can serve as reliable early indicators of the final generated response's information coverage. Through experiments across two text RAG benchmarks (TREC NeuCLIR 2024 and TREC RAG 2024) and one multimodal benchmark (WikiVideo), we analyze 15 text retrieval stacks and 10 multimodal retrieval stacks across four RAG pipelines and multiple evaluation frameworks (Auto-ARGUE and MiRAGE). Our findings demonstrate strong correlations between coverage-based retrieval metrics and nugget coverage in generated responses at both topic and system levels. This relationship holds most strongly when retrieval objectives align with generation goals, though more complex iterative RAG pipelines can partially decouple generation quality from retrieval effectiveness. These findings provide empirical support for using retrieval metrics as proxies for RAG performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。