揭示检索结果呈现方式如何影响阅读器使用支持证据的效果
Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support

- 提出'证据接口'概念,分析检索结果格式对阅读器的影响
- 发现顶级检索窗口的性能损失多由关键支持信息缺失导致
- 建议在评估时同时报告答案得分与完整支持覆盖度
在多跳RAG评估中,仅看top-k答案得分可能掩盖两种不同失败:检索窗口可能遗漏支持链的一部分,或包含阅读器难以利用的形式。我们称这种面向阅读器的检索证据呈现方式为‘证据接口’。通过三个带有支持标注的多跳问答基准,对比使用原始上下文、检索窗口和黄金支持诊断渲染的适配阅读器。结果区分了支持可用性问题与阅读器-接口交互效应。只有在检查完整标注支持链是否留存后,top-k窗口才具有可解释性:若支持链完整,短排名窗口可媲美甚至优于原始上下文;若不完整,则缺失支持解释了大部分性能下降。黄金支持先行训练能提升适配阅读器表现;在2Wiki和MuSiQue上,支持监督排序器提升了覆盖率,并以更低提示成本恢复原始上下文质量,同时保留黄金上限。支持移除测试进一步表明,性能提升源于暴露的证据,而不仅是答案先验。因此,在支持标注评估中,应同时报告top-k答案得分与完整支持覆盖度。
原文摘要 · Abstract (English)
In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain, or it may contain support in a form the adapted reader does not use well. We call this reader-facing form of retrieved evidence an evidence interface. Using three support-annotated multi-hop QA benchmarks, we compare matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings. These comparisons distinguish support-availability failures from remaining reader-interface effects. Top-k windows become interpretable only after checking whether the complete annotated support chain survives: when it does, short ranked windows can match or improve over raw context; when it does not, missing support explains much of the loss. Gold support-first improves matched readers; on 2Wiki and MuSiQue, a support-supervised ranker raises coverage and recovers raw-context quality at lower prompt cost, while retaining gold headroom. Support-removal checks further show that the gains rely on exposed evidence, not only answer priors. On support-annotated evaluations, top-k answer scores should therefore be reported together with complete-support coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。