arXiv:2509.15577cs.CLcs.AI2025-09ACL被引 1

让检索文本更实用:通过推理过程监督重写,提升问答效果

Relevance to Utility: Process-Supervised Rewrite for RAG

  • 通过推理时的重写与回答联合观察,捕捉文档真实价值
  • 在多个开放域问答数据集上,性能超越强基线模型
  • 适合关注RAG系统中检索-生成对齐问题的研究者

检索增强生成系统常面临优化检索相关性与生成实用性之间的差距。即使检索到的文档主题相关,也可能缺乏有效推理所需内容。现有桥接模块尝试重写检索文本以改善生成效果,但我们发现它们未能捕捉‘文档实用性’。本文提出R2U,其核心在于通过联合观察重写与回答过程来近似真实实用性。为提升蒸馏可靠性,我们进一步通过测量生成器在重写上下文下的答案提升程度,构建实用性改进的监督信号,用于微调与偏好优化。我们在多个开放域问答基准上评估该方法,实验证明其持续优于强基线模型。

原文摘要 · Abstract (English)

Retrieval-augmented generation systems often suffer from a gap between optimizing retrieval relevance and generative utility. With such a gap, retrieved documents may be topically relevant but still lack the content needed for effective reasoning during generation. While existing bridge modules attempt to rewrite the retrieved text for better generation, we show how they fail by not capturing "document utility". In this work, we propose R2U, with a key distinction of approximating true utility through joint observation of rewriting and answering in the reasoning process. To distill, R2U scale such supervision to enhance reliability in distillation. We further construct utility-improvement supervision by measuring the generator's gain of the answer under the rewritten context, yielding signals for fine-tuning and preference optimization. We evaluate our method across multiple open-domain question-answering benchmarks. The empirical results demonstrate consistent improvements over strong bridging baselines

RAG生成优化检索增强实用度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。