用问题分解引导检索,让多模态模型答得更准更可靠。
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation

- 通过拆解问题并引导检索,提升知识获取的准确性。
- 在三个基准上均显著优于现有方法,最高提升6.2%。
- 适合需要复杂推理和跨领域知识的视觉问答任务。
随着多模态研究与深度学习的发展,多模态大语言模型(MLLMs)已成为视觉-语言任务的重要范式。作为视觉-语言研究的核心问题,视觉问答(VQA)在开放域场景中日益依赖外部知识,因此采用MLLMs提升性能。本文提出一种逻辑提示策略,将思维链(CoT)推理与视觉问题分解(VQD)融合,称为CoVQD,以更精准地引导检索过程。基于此,我们构建了CoVQD引导的检索增强生成框架(CgRAG),使MLLM能获取更全面且连贯的外部知识,并受益于结构化视觉-文本推理指导,从而在复杂跨域VQA场景中提升泛化能力与可靠性。在E-VQA、InfoSeek和OKVQA等基准上的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of multimodal tasks. As a core problem in vision-language research, Visual Question Answering (VQA) has increasingly employed MLLMs to improve performance, particularly in open-domain settings where external knowledge is essential. In this work, we aim to further enhance retrieval-based VQA by more effectively integrating MLLMs with structured reasoning and knowledge acquisition. We introduce a logical prompting strategy that fuses Chain-of-Thought (CoT) reasoning with Visual Question Decomposition (VQD), termed CoVQD, to guide retrieval toward more accurate and relevant knowledge for MLLM inference. Building on this idea, we propose a new framework, CoVQD-guided RAG (CgRAG), which enables MLLMs to access more comprehensive and coherent external knowledge while benefiting from structured visual-text reasoning guidance, thereby improving generalization and reliability in complex cross-domain VQA scenarios. Extensive experiments on E-VQA, InfoSeek, and OKVQA benchmarks demonstrate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。