arXiv:2505.10089cs.CL2025-05EMNLP被引 17

评测大模型跨语言检索生成能力,发现语言匹配与跨语言推理双重挑战

XRAG: Cross-lingual Retrieval-Augmented Generation

  • 基于新闻构建跨语言检索生成数据集,要求复杂推理
  • 模型在单语检索中常答错语言,多语检索中难跨语言推理
  • 适合研究大模型跨语言理解与推理能力的研究者使用

我们提出XRAG,一个新型基准,用于评估大模型在用户语言与检索结果语言不一致的跨语言检索增强生成(RAG)场景中的生成能力。该数据集源自近期新闻文章,确保问题需外部知识才能回答。涵盖单语和多语检索的真实场景,并为每个检索文档提供相关性标注。新颖的数据集构建流程使问题需复杂推理,表现为人类与大模型性能差距显著。因此,即使不考虑跨语言复杂性,XRAG也已成为研究大模型推理能力的重要基准。对五种大模型的实验揭示了跨语言RAG中两个此前未报告的挑战:1)在单语检索设置下,所有模型均难以保证回答的语言正确性;2)在多语检索设置下,主要难点在于跨语言信息推理,而非生成非英文文本。

原文摘要 · Abstract (English)

We propose XRAG, a novel benchmark designed to evaluate the generation abilities of LLMs in cross-lingual Retrieval-Augmented Generation (RAG) settings where the user language does not match the retrieval results. XRAG is constructed from recent news articles to ensure that its questions require external knowledge to be answered. It covers the real-world scenarios of monolingual and multilingual retrieval, and provides relevancy annotations for each retrieved document. Our novel dataset construction pipeline results in questions that require complex reasoning, as evidenced by the significant gap between human and LLM performance. Consequently, XRAG serves as a valuable benchmark for studying LLM reasoning abilities, even before considering the additional cross-lingual complexity. Experimental results on five LLMs uncover two previously unreported challenges in cross-lingual RAG: 1) in the monolingual retrieval setting, all evaluated models struggle with response language correctness; 2) in the multilingual retrieval setting, the main challenge lies in reasoning over retrieved information across languages rather than generation of non-English text.

跨语言检索生成大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。