arXiv:2507.23242cs.CVcs.CL2025-07被引 3

无需人工标注,用可验证搜索奖励优化检索增强生成的查询重写。

Annotation-Free Reinforcement Learning Query Rewriting via Verifiable Search Reward

  • 基于可验证搜索奖励的强化学习框架,自动优化查询重写。
  • 在视觉文档上检索性能提升最高达3.9倍,语义检索提升3.5倍。
  • 适用于多模态与不同索引领域,适合工业级RAG系统部署。

优化检索增强生成(RAG)系统的查询面临重大挑战,尤其在跨多种模态索引时。我们提出RL-QR,一种无需人工标注的强化学习查询重写框架,摆脱了对昂贵人工标注数据的依赖。通过利用与索引对齐的合成查询生成的可验证搜索奖励,RL-QR克服了人工标注依赖,扩展至多种模态和索引领域。实验表明该框架具有强鲁棒性,在MTEB VIDORE V2基准上,对词汇检索器的检索性能提升高达3.9倍,对语义检索器提升达3.5倍;在MS MARCO v2.1及内部工业数据集上均实现5%至10%的稳定提升。

原文摘要 · Abstract (English)

Optimizing queries for Retrieval-Augmented Generation (RAG) systems poses a significant challenge, particularly across diverse modal indices. We introduce RL-QR, a novel annotation-free reinforcement learning framework for query rewriting that eliminates the need for costly human-annotated data. By leveraging verifiable search rewards derived from index-aligned synthetic queries, RL-QR overcomes human-annotation dependencies, extending its applicability to various modalities and index domains. Experimental results demonstrate the framework's robustness, achieving substantial retrieval performance gains of up to 3.9$\times$ on lexical retrievers and 3.5$\times$ on semantic retrievers on the MTEB VIDORE V2 benchmark for unstructured visual documents, along with consistent 5\% to 10\% improvements on MS MARCO v2.1 and internal industrial datasets.

RAG强化学习查询重写无标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。