构建卢森堡语文档跨语言图文检索基准,验证图像检索优于文本检索。
LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

- 以卢森堡语文档图像为索引,支持多语言查询的跨语言检索评测
- 图像检索模型在所有语言上表现均优于传统文本检索方法
- 多语言微调中加入卢森堡语显著提升本地查询效果,适合小语种研究者
近年来,如ColPali等页面图像检索器在视觉丰富的文档检索中取得进展,但在跨语言、低资源场景下的表现仍不明确。本文提出LëtzCross,一个面向卢森堡语PDF文档的跨语言页面级检索基准,将文档页面作为图像索引,查询语言包括英语、法语、德语和卢森堡语。该基准融合文本导向与视觉对齐的问答对,覆盖基于PDF的RAG中的文本与视觉检索需求。通过LëtzCross对比基于OCR的文本检索器与ColPali风格的页面图像检索器,发现后者在所有查询语言下表现更优。此外,评估了单语言与多语言微调策略:单语言微调中,法语设置在卢森堡语查询上达到最高平均性能;多语言微调中,包含卢森堡语时结果最佳,显著提升卢森堡语查询的检索效果。
原文摘要 · Abstract (English)
Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-lingual, low-resource settings. We introduce LëtzCross, a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, with document pages indexed as images and queries provided in English, French, German, and Luxembourgish. The benchmark combines text-focused QA pairs with visually grounded QA pairs, covering both textual and visual retrieval needs in PDF-based RAG. We use LëtzCross to compare OCR-based text-only retrievers with ColPali-style page-image retrievers and find that the latter perform better across query languages in this system-level comparison. We also examine single-language and multilingual fine-tuning. Fine-tuning transfers across query languages, with French yielding the highest mean performance on Luxembourgish queries among the single-language settings. In the multilingual setting, including Luxembourgish gives the strongest results and substantially improves retrieval for Luxembourgish queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。