arXiv:2605.30027cs.CVcs.IR2026-05KDD

提升多模态文档检索效果,兼顾布局信息与少样本泛化能力。

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

论文配图:DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
图 1 · 摘自论文原文
  • 用布局感知的稀疏嵌入增强视觉编码,无需OCR即可融合结构信息。
  • 引入基于推理增强演示的可泛化重排序器,在少样本下准确率显著提升。
  • 构建新基准MultiDocR,支持多维度评估,推动领域可靠评测发展。

多模态文档包含表格、图表和版式等多种元素,使检索任务复杂化。现有方法通常结合密集视觉嵌入模型与监督重排序器以实现高精度检索,但存在固有局限:首先,密集嵌入的粗粒度特性易掩盖显式语义,难以利用结构关键信息;其次,监督重排序模型受制于泛化瓶颈,性能高度依赖特定领域训练数据。此外,现有基准常缺乏多样评估维度和全面的相关性标注,限制了可靠评估。为此,我们提出DocRetriever,一个即插即用框架。通过布局感知的稀疏嵌入技术增强视觉检索,实现无需光学字符识别(OCR)的有效混合编码。同时引入可泛化的重排序器,借助推理增强演示与优化采样策略,提升少样本场景下的准确性。最后,我们构建新基准MultiDocR,支持更严格的评估。跨多个基准的实验验证了DocRetriever在性能上优于当前最先进方法。

原文摘要 · Abstract (English)

Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typically combine dense visual embedding models with supervised rerankers to achieve high-precision retrieval, they face inherent limitations. First, the coarse-grained nature of dense embeddings tends to obfuscate explicit semantics, failing to leverage structurally salient information. Second, supervised reranking models suffer from generalization bottlenecks, as their performance heavily relies on domain-specific training data. Furthermore, existing benchmarks often lack diverse assessment dimensions and comprehensive relevance annotations, limiting reliable evaluation. To address these challenges, we propose DocRetriever, a plug-and-play framework. It enhances visual retrieval via a layout-aware sparse embedding technique, enabling effective hybrid encoding without the overhead of optical character recognition (OCR). We also introduce a generalizable reranker that leverages reasoning-augmented demonstrations and optimized sampling to improve accuracy in few-shot settings. Finally, we construct a new benchmark, MultiDocR, to enable more rigorous evaluation. Experiments across diverse benchmarks validate DocRetriever's superiority over state-of-the-art methods.

文档检索多模态稀疏嵌入少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。