arXiv:2604.25182cs.CLcs.IR2026-04中稿 · SIGIR 2026被引 1

跨语言检索增强生成新框架,让多语种知识更好互补

CroSearch-R1: Better Leveraging Cross-lingual Knowledge for Retrieval-Augmented Generation

论文配图:CroSearch-R1: Better Leveraging Cross-lingual Knowledge for Retrieval-Augmented Generation
图 1 · 摘自论文原文
  • 用强化学习动态对齐多语言知识到统一空间
  • 在多个语言数据集上提升事实准确率,显著优于基线
  • 适合需要多语种知识融合的生成系统开发者

多语言语料库可能包含可补充和修正原始语言事实的其他语言知识。然而,简单拼接不同语言的知识片段至上下文可能因语言差异导致效果下降。为更好利用多语言知识,我们提出CroSearch-R1,一种基于搜索增强的强化学习框架,将多语言知识融入组相对策略优化(GRPO)过程。该方法采用多轮检索策略,结合跨语言知识整合,动态将其他语言的知识作为补充证据对齐至统一表示空间;同时引入多语言回放机制,优化跨语言推理迁移能力。实验表明,该框架有效利用跨语言互补性,在多语言语料上显著提升RAG的效果。

原文摘要 · Abstract (English)

A multilingual collection may contain useful knowledge in other languages to supplement and correct the facts in the original language for Retrieval-Augmented Generation (RAG). However, the vanilla approach that simply concatenates multiple pieces of knowledge from different languages into the context may fail to improve effectiveness due to the potential disparities across languages. To better leverage multilingual knowledge, we propose CroSearch-R1, a search-augmented reinforcement learning framework to integrate multilingual knowledge into the Group Relative Policy Optimization (GRPO) process. In particular, the approach adopts a multi-turn retrieval strategy with cross-lingual knowledge integration to dynamically align the knowledge from other languages as supplementary evidence into a unified representation space. Furthermore, we introduce a multilingual rollout mechanism to optimize reasoning transferability across languages. Experimental results demonstrate that our framework effectively leverages cross-lingual complementarity and improves the effectiveness of RAG with multilingual collections.

检索增强多语言强化学习知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。