arXiv:2603.13349cs.CVcs.AI2026-03

用多尺度视觉编码提升文档检索效率与精度

MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval

  • 分层多分辨率编码结合视觉语言模型,融合不同尺度视觉特征
  • 在仅用50%视觉标记量时,性能超越ColPali
  • 适合需要高效文档检索的场景,如大规模电子文档系统

视觉文档检索(VDR)需要同时捕捉细粒度视觉细节和全局文档结构,以保证检索效果并维持计算效率。现有模型在处理高分辨率文档时难以平衡效果与效率:要么丢失细节信息,要么生成过多视觉标记,导致索引开销大、检索延迟高。本文重新思考视觉编码机制,提出X-VisEmb范式,涵盖多分辨率采样与编码、跨粒度特征融合及自适应表示压缩。初步研究表明该方法能有效捕获多尺度互补视觉线索。基于此,我们构建MURE框架,利用视觉语言模型作为分层多分辨率编码器,引入层次化马特里什卡表征学习(RMRL)实现高效特征融合,并采用语义感知的分层聚类机制压缩视觉标记。在两个主流VDR基准上的实验表明,MURE持续优于强基线,且仅用50%的视觉标记预算即超越ColPali。

原文摘要 · Abstract (English)

Visual Document Retrieval (VDR) requires representations that capture both fine-grained visual details and global document structure to ensure retrieval efficacy while maintaining computational efficiency. Existing VDR models struggle to balance effectiveness and efficiency when processing high-resolution documents: they often either lose fine-grained information or generate an excessive number of visual tokens, resulting in significant indexing overhead and high retrieval latency. In this work, we rethink the visual encoding mechanism and propose a new X-VisEmb paradigm that progresses from multi-resolution sampling and encoding, through cross-granularity feature fusion, to adaptive representation distillation. A preliminary study validates its feasibility and effectiveness in capturing complementary visual cues at varying scales. Building on the insights, we develop MURE, a novel framework that employs VLMs as a hierarchical multi-resolution encoder, integrates resolution-level Matryoshka representation learning (RMRL) for effective feature fusion, and applies a semantic-aware hierarchical clustering mechanism for visual token compression. Experiments on two widely used VDR benchmarks show that our MURE framework consistently beats strong baselines. Furthermore, it significantly outperforms ColPali with only 50% of its visual token budget.

文档检索多分辨率视觉语言模型表征压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。