MIRAGE通过分层分解提升图像检索精度与效率,减少冗余计算。
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
- 采用多粒度分层结构增强查询与图像对象对齐。
- 利用跨层级相似性一致性降低冗余匹配,计算量减少3.5倍。
- 自动配置参数适配不同数据集,实用性强。
为有效利用用户特定数据,多模态大语言模型(MLLM)应用中采用检索增强生成(RAG)。然而,传统检索方法常受限于精度不足。近期的多向量检索(MVR)通过分解查询并匹配分割图像提升精度,但仍存在准确率与效率不优的问题,忽视了查询与不同图像对象间的对齐关系,以及细粒度图像片段的冗余。本文提出一种高效的图像检索调度框架MIRAGE。首先,引入新颖的分层范式,使用多种中间粒度处理不同图像对象以增强对齐。其次,通过跨层级相似性一致性与层级稀疏性最小化冗余检索,减少不必要的匹配计算。此外,针对不同数据集自动配置参数以提升实际适用性。实验表明,MIRAGE不仅显著提升准确率,且相比现有MVR系统计算量减少高达3.5倍。
原文摘要 · Abstract (English)
To effectively leverage user-specific data, retrieval augmented generation (RAG) is employed in multimodal large language model (MLLM) applications. However, conventional retrieval approaches often suffer from limited retrieval accuracy. Recent advances in multi-vector retrieval (MVR) improve accuracy by decomposing queries and matching against segmented images. They still suffer from sub-optimal accuracy and efficiency, overlooking alignment between the query and varying image objects and redundant fine-grained image segments. In this work, we present an efficient scheduling framework for image retrieval - MIRAGE. First, we introduce a novel hierarchical paradigm, employing multiple intermediate granularities for varying image objects to enhance alignment. Second, we minimize redundancy in retrieval by leveraging cross-hierarchy similarity consistency and hierarchy sparsity to minimize unnecessary matching computation. Furthermore, we configure parameters for each dataset automatically for practicality across diverse scenarios. Our empirical study shows that, MIRAGE not only achieves substantial accuracy improvements but also reduces computation by up to 3.5 times over the existing MVR system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。