arXiv:2603.04238cs.CL2026-03

改进文档表示比检索机制更能提升多语言视觉RAG性能

Retrieval or Representation? Reassessing Benchmark Gaps in Multilingual and Visually Rich RAG

  • 固定检索方法,只改变文本转录与预处理方式
  • BM25在多语言和视觉基准上可大幅缩小性能差距
  • 适合关注模型真实进步来源的研究者

检索增强生成(RAG)通过外部文档和实时信息来增强语言模型。传统检索系统依赖词法方法如BM25,根据术语重叠和语料库权重排序文档。端到端多模态检索器在大规模查询-文档数据集上训练,声称在具有复杂视觉布局的多语言文档上显著优于传统方法。我们证明,文档表示的改进是基准性能提升的主要原因。在保持检索机制不变的前提下,系统性地调整文本转录与预处理方法后,发现BM25可在多语言和视觉基准上恢复巨大性能差距。研究呼吁建立解耦的评估基准,分别衡量文本转录与检索能力,使领域能正确归因进展,并聚焦真正关键问题。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) is a common way to ground language models in external documents and up-to-date information. Classical retrieval systems relied on lexical methods such as BM25, which rank documents by term overlap with corpus-level weighting. End-to-end multimodal retrievers trained on large query-document datasets claim substantial improvements over these approaches, especially for multilingual documents with complex visual layouts. We demonstrate that better document representation is the primary driver of benchmark improvements. By systematically varying transcription and preprocessing methods while holding the retrieval mechanism fixed, we demonstrate that BM25 can recover large gaps on multilingual and visual benchmarks. Our findings call for decomposed evaluation benchmarks that separately measure transcription and retrieval capabilities, enabling the field to correctly attribute progress and focus effort where it matters.

RAG多语言视觉检索评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。