通过多粒度多模态协同检索,提升视觉问答系统对知识的精准获取能力。
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
- 采用从粗到细的多阶段检索,融合不同粒度与模态信息
- 在InfoSeek和Encyclopedic-VQA上达到最新最好性能
- 适合需要高效跨模态知识检索的研究与应用
视觉-语言检索增强生成(RAG)已成为解决基于知识的视觉问答(KB-VQA)的有效方法,该任务需获取图像内容之外的外部知识。RAG系统的有效性取决于多模态检索,而查询与知识库中存在多样化的模态和知识粒度,导致检索本身极具挑战性。现有方法尚未充分挖掘这些元素间的协同潜力。本文提出一种新型多模态RAG系统,采用从粗到细的多步检索策略,协调多种粒度与模态以提升检索效率。系统首先进行宽范围初始搜索,对齐知识粒度以支持跨模态检索;随后通过多模态融合重排序,捕捉细微的多模态信息以选出最优实体;最后由文本重排序器筛选最相关细粒度段落用于生成增强。在InfoSeek和Encyclopedic-VQA基准上的大量实验表明,本方法在检索性能上达到当前最佳水平,并取得极具竞争力的问答结果,验证了其在推进KB-VQA系统方面的有效性。
原文摘要 · Abstract (English)
Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented in images. The effectiveness of Vision-language RAG systems hinges on multimodal retrieval, which is inherently challenging due to the diverse modalities and knowledge granularities in both queries and knowledge bases. Existing methods have not fully tapped into the potential interplay between these elements. We propose a multimodal RAG system featuring a coarse-to-fine, multi-step retrieval that harmonizes multiple granularities and modalities to enhance efficacy. Our system begins with a broad initial search aligning knowledge granularity for cross-modal retrieval, followed by a multimodal fusion reranking to capture the nuanced multimodal information for top entity selection. A text reranker then filters out the most relevant fine-grained section for augmented generation. Extensive experiments on the InfoSeek and Encyclopedic-VQA benchmarks show our method achieves state-of-the-art retrieval performance and highly competitive answering results, underscoring its effectiveness in advancing KB-VQA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。