构建多语言视觉文档检索基准,评估模型跨语言图文理解能力。
MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark
- 基于18种语言的MIRACL数据集扩展,生成高质量多语言视觉检索问题。
- 通过移除简单负样本压缩数据集,提升计算效率同时保持挑战性。
- 发现视觉模型在多语言任务上表现显著落后于文本模型,差距达59.7%。
文档检索是搜索与检索增强生成(RAG)应用中的关键任务。尽管大语言模型提升了文本检索的准确性,但包含表格、图表等复杂布局和视觉元素的文档难以通过纯文本表示。近年来,基于图像的文档检索流程兴起,利用视觉大模型(VLMs)根据查询检索相关页面图像。然而,现有视觉文档检索评估基准存在局限:仅覆盖英语、依赖合成问题、语料规模小。为此,我们提出MIRACL-VISION,一个支持18种语言的多语言视觉文档检索评估基准。该基准基于人工标注的MIRACL数据集扩展而来,确保问题质量。为降低计算开销,我们设计方法剔除“简单负样本”,在保持难度的同时缩减语料规模。我们使用主流公开文本与图像模型进行了广泛实验,结果表明当前最先进的基于视觉大模型的嵌入模型在多语言能力上存在明显不足,其检索准确率相比文本模型低至59.7%;即使在英语上,视觉模型仍比文本模型低12.1%。MIRACL-VISION是一个具有挑战性、代表性强的多语言视觉检索评估基准,有助于推动社区构建更鲁棒的文档检索模型。
原文摘要 · Abstract (English)
Document retrieval is an important task for search and Retrieval-Augmented Generation (RAG) applications. Large Language Models (LLMs) have contributed to improving the accuracy of text-based document retrieval. However, documents with complex layout and visual elements like tables, charts and infographics are not perfectly represented in textual format. Recently, image-based document retrieval pipelines have become popular, which use visual large language models (VLMs) to retrieve relevant page images given a query. Current evaluation benchmarks on visual document retrieval are limited, as they primarily focus only English language, rely on synthetically generated questions and offer a small corpus size. Therefore, we introduce MIRACL-VISION, a multilingual visual document retrieval evaluation benchmark. MIRACL-VISION covers 18 languages, and is an extension of the MIRACL dataset, a popular benchmark to evaluate text-based multilingual retrieval pipelines. MIRACL was built using a human-intensive annotation process to generate high-quality questions. In order to reduce MIRACL-VISION corpus size to make evaluation more compute friendly while keeping the datasets challenging, we have designed a method for eliminating the "easy" negatives from the corpus. We conducted extensive experiments comparing MIRACL-VISION with other benchmarks, using popular public text and image models. We observe a gap in state-of-the-art VLM-based embedding models on multilingual capabilities, with up to 59.7% lower retrieval accuracy than a text-based retrieval models. Even for the English language, the visual models retrieval accuracy is 12.1% lower compared to text-based models. MIRACL-VISION is a challenging, representative, multilingual evaluation benchmark for visual retrieval pipelines and will help the community build robust models for document retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。