arXiv:2603.01666cs.CLcs.IR2026-03被引 5

用布局解析生成紧凑多向量,存储降95%还更准

Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations

  • 先解析文档布局,生成少量结构化子图嵌入
  • 存储减少95%以上,跨多个基准性能提升显著
  • 适合需要高效部署的视觉文档检索系统

充分挖掘视觉丰富文档的潜力,需具备理解文本与复杂布局能力的检索系统,这是视觉文档检索(VDR)的核心挑战。当前主流多向量架构虽强大,却面临严重存储瓶颈,现有优化策略如嵌入合并、剪枝或使用抽象标记,无法在不损害性能或忽略关键布局线索的前提下解决该问题。为此,我们提出ColParse,一种新范式:利用文档解析模型生成少量布局感知的子图像嵌入,并将其与全局页面向量融合,构建紧凑且结构敏感的多向量表示。大量实验证明,该方法在保持高性能的同时,将存储需求降低超过95%,并在多个基准和基础模型上实现显著性能提升。ColParse成功弥合了多向量检索的细粒度准确性与大规模部署的实际需求之间的关键鸿沟,为高效、可解释的多模态信息系统的构建开辟新路径。

原文摘要 · Abstract (English)

Harnessing the full potential of visually-rich documents requires retrieval systems that understand not just text, but intricate layouts, a core challenge in Visual Document Retrieval (VDR). The prevailing multi-vector architectures, while powerful, face a crucial storage bottleneck that current optimization strategies, such as embedding merging, pruning, or using abstract tokens, fail to resolve without compromising performance or ignoring vital layout cues. To address this, we introduce ColParse, a novel paradigm that leverages a document parsing model to generate a small set of layout-informed sub-image embeddings, which are then fused with a global page-level vector to create a compact and structurally-aware multi-vector representation. Extensive experiments demonstrate that our method reduces storage requirements by over 95% while simultaneously yielding significant performance gains across numerous benchmarks and base models. ColParse thus bridges the critical gap between the fine-grained accuracy of multi-vector retrieval and the practical demands of large-scale deployment, offering a new path towards efficient and interpretable multimodal information systems.

文档检索多向量布局解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。