让文档检索模型按需调整计算量,又快又省。
MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework

- 用2D嵌套训练框架,同时调节模型深度和向量数量
- 比直接截断保留更高精度,存储与计算开销大幅下降
- 适合需要灵活部署的文档检索系统
多向量视觉文档检索器通过深度视觉语言模型为每页生成多个向量,实现精细匹配,但带来高昂的存储与计算成本。现有效率优化通常只针对部分预算,缺乏统一方法来同时平衡向量宽度与编码器深度。为此,我们提出MM-Matryoshka,一种支持2D可调预算的多模态嵌套训练框架,使ColPali类多向量检索在推理时可灵活选择不同预算组合,无需为每种配置训练独立模型。在多个主流骨干网络上的实验表明,该方法在显著降低存储与计算开销的同时,仍保持远高于直接截断基线的检索质量,为高效视觉文档检索提供了鲁棒的预算弹性能力。
原文摘要 · Abstract (English)
Multi-vector visual document retrievers achieve strong fine-grained matching by representing each page with multiple vectors from deep Vision-Language Models (VLMs), but this design makes deployment expensive in both storage and computational overhead. Existing efficiency techniques usually optimize only part of this budget, leaving multimodal retrievers without a unified way to trade accuracy for both vector width and encoder depth. Therefore, we propose MM-Matryoshka, a 2D Matryoshka training framework for budget-elastic Visual Document Retrieval (VDR), enabling ColPali-style multi-vector retrieval elastic along both dimension and layer. At inference time, a single retriever can select a 2D selectable budget without training separate models for different budgets. Through comprehensive experiments across multiple representative backbones, we demonstrate that by retaining significantly higher quality than direct truncation baselines while substantially reducing storage and computational overhead, MM-Matryoshka can offer robust budget elasticity for efficient VDR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。