提出剪枝-合并框架,高效压缩视觉文档检索向量。
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
- 先自适应剪枝低信息块,再分层合并高信号向量。
- 在29个数据集上实现近无损压缩,高压缩比下仍保持强性能。
- 适合需要高效检索海量图文文档的场景。
视觉文档检索(VDR)旨在从海量视觉丰富文档中检索相关页面,对当前多模态检索应用具有重要意义。现有先进多向量范式虽性能优异,但存在高昂开销问题,现有效率方法如剪枝与合并难以兼顾压缩率与特征保真度,导致压缩与精度间的权衡困境。为此,我们提出全新的两阶段剪枝-合并框架,协同互补两类方法。首先通过自适应剪枝阶段过滤低信息图像块,生成高信噪比嵌入集合;随后在该预筛选基础上进行层次化合并,有效总结语义内容,避免单阶段方法因噪声导致的特征稀释。在29个VDR数据集上的大量实验表明,本框架持续优于现有方法,显著拓展了近无损压缩范围,并在高压缩比下仍具备鲁棒性能。
原文摘要 · Abstract (English)
Visual Document Retrieval (VDR), which aims to retrieve relevant pages within vast corpora of visually-rich documents, is of significance in current multimodal retrieval applications. The state-of-the-art multi-vector paradigm excels in performance but suffers from prohibitive overhead, a problem that current efficiency methods like pruning and merging address imperfectly, creating a difficult trade-off between compression rate and feature fidelity. To overcome this dilemma, we introduce Prune-then-Merge, a novel two-stage framework that synergizes these complementary approaches. Our method first employs an adaptive pruning stage to filter out low-information patches, creating a refined, high-signal set of embeddings. Subsequently, a hierarchical merging stage compresses this pre-filtered set, effectively summarizing semantic content without the noise-induced feature dilution seen in single-stage methods. Extensive experiments on 29 VDR datasets demonstrate that our framework consistently outperforms existing methods, significantly extending the near-lossless compression range and providing robust performance at high compression ratios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。