arXiv:2606.04240cs.CVcs.AI2026-06综述

构建能处理图文混排文档的统一检索系统,挑战多模态信息融合新高度。

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1)

  • 基于Qwen2-VL的多模态大模型嵌入,融合文本与图像特征
  • 在两项任务上均取得高召回率,最佳系统达Recall@5=0.73
  • 无需微调的融合方案表现逼近最优微调系统

面向图文混排文档的检索对多模态检索增强生成至关重要,但多数检索器仍忽略视觉信息。本次挑战赛(EReL@MIR 2025)要求参赛者构建单一系统,同时应对两类任务:基于文本查询的长文档内页面闭集检索(MMDocIR),以及基于图像或图文组合查询的开放域百科式段落检索(M2KR)。系统按两项任务的平均召回率(Recall@1,3,5)宏平均排名。共吸引22支队伍、586次提交。最终前三名均采用基于Qwen2-VL的解码器架构多模态大模型嵌入,而非传统CLIP式编码器;其差异在于:是否微调集成、训练自由的多路径融合搭配强视觉-语言重排序器,或零样本后期交互。无训练系统仅落后于微调优胜者0.1分。

原文摘要 · Abstract (English)

Retrieval over visually-rich documents, pages that interleave text with figures, tables, and charts, is essential for multimodal retrieval-augmented generation, yet most retrievers still discard the visual channel. The \emph{Multimodal Document Retrieval Challenge}, Track~1 of the MIR Challenge at the first EReL@MIR workshop, co-located with The Web Conference 2025, asks participants to build a \emph{single} retrieval system that handles two complementary regimes: closed-set document page retrieval within long documents from a text query (MMDocIR), and open-domain retrieval of Wikipedia-style passages from an image or image-plus-text query (M2KR). Systems are ranked by the macro-average of mean Recall@$\{1,3,5\}$ over the two tasks. The challenge drew 455 entrants and 586 submissions across 22 teams. This report describes the challenge design, datasets, and evaluation protocol; reports the final standings; and analyses the three winning teams' systems. All three build on decoder-based Multimodal-LLM embedders from the Qwen2-VL family rather than on CLIP-style encoders, and differ chiefly in whether they reach the top through fine-tuned ensembles, training-free multi-route fusion with a strong vision-language re-ranker, or zero-shot late interaction. The training-free system finished within $0.1$ point of the fine-tuned winner.

多模态检索文档理解大模型应用图文混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。