研究检索结果是否含支持生成的证据,发现其对下游任务的帮助不一致。
Assessing the Downstream Utility of Evidence-Aware Retrieval in RAG

- 用答案支持信号评估检索效果,而非仅看主题相关性。
- 该信号能改变检索排序,但不能稳定提升生成质量或训练效果。
- 适合用于人工筛选证据,但对自动系统决策帮助有限。
针对检索增强生成(RAG)的检索评估正从单纯关注主题相关性转向考察检索段落是否包含支持生成的答案证据。本文在五个检索基准和一个端到端的TREC RAG 2025设置下,考察答案支持信号在四种角色中的表现:比较不同检索器、指导检索训练与系统选择、预测下游答案质量、以及过滤提供给生成器的证据。结果显示,该信号虽改变检索排序,但其下游价值并不一致:无法可靠提升检索器训练效果;系统选择的收益依赖于生成器使用证据的方式;基于该信号的检索分数也无法稳健预测未见话题下的答案质量。人工干预实验表明,优先保留含有效证据的段落确实可行,但不同评估者对最终答案质量改善的判断存在分歧。这说明,使检索评估更贴近生成所需证据,并不必然提升所有下游应用的可靠性。因此,RAG评估方法应根据其具体用途进行针对性评估。
原文摘要 · Abstract (English)
Retrieval evaluation for retrieval-augmented generation (RAG) is increasingly designed around whether retrieved passages contain evidence that can support generation, rather than topical relevance alone. We study whether this closer alignment with downstream evidence needs also makes retrieval evaluation more useful for the decisions built from it. Across five retrieval benchmarks and an end-to-end TREC RAG 2025 setting, we examine an answer-support signal in four roles: comparing retrievers, guiding retrieval training and system selection, predicting downstream answer quality, and filtering the evidence supplied to a generator. The signal changes retrieval rankings, but its downstream value is not uniform. It does not reliably improve retriever training; the benefit of using it for system selection depends on how the generator is instructed to use the retrieved evidence; and retrieval scores based on it do not robustly predict answer quality on unseen topics. In a direct evidence intervention, human annotators confirm that filtering preferentially preserves passages containing useful answer evidence, yet different answer evaluators reach different conclusions about whether the resulting answers improve. These results show that making retrieval evaluation more closely reflect the evidence needed for generation does not by itself make every downstream use of that evaluation more reliable. RAG evaluation methods should therefore be assessed with respect to the particular comparisons, decisions, and conclusions they are intended to support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。