arXiv:2502.20964cs.CVcs.AI2025-02被引 2

通过细粒度知识结构提升视觉问答的精准度与推理能力

Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering

  • 将多模态数据片段结构化为细粒度知识单元,增强信息组织性
  • 在四个基准上平均提升3%,最高达11%的准确率
  • 适合需要高精度视觉推理与外部知识融合的研究者

视觉问答(VQA)旨在利用图像信息回答自然语言问题。尽管先进的多模态大模型(如GPT-4o)在VQA任务上表现优异,但往往难以获取特定领域或最新知识。为此,基于外部知识库的检索增强生成(RAG),即KB-VQA,成为有前景的解决方案。然而,传统单模态检索方法通过图像转文本描述,常导致关键视觉细节丢失。本文提出两项创新:首先,引入由多模态数据片段(如文本片段、实体图像等)构成的细粒度知识单元,并以结构化方式组织管理,使知识结构本身提升检索质量;其次,提出知识单元检索增强生成框架(KU-RAG),无缝融合细粒度检索与多模态大模型。该框架不仅实现精准知识检索,还通过知识修正链增强推理能力。实验表明,本方法在四个基准上持续优于现有KB-VQA方法,平均提升约3%,最佳情况下达11%。

原文摘要 · Abstract (English)

Visual Question Answering (VQA) focuses on providing answers to natural language questions by utilizing information from images. Although cutting-edge multimodal large language models (MLLMs) such as GPT-4o achieve strong performance on VQA tasks, they frequently fall short in accessing domain-specific or the latest knowledge. To mitigate this issue, retrieval-augmented generation (RAG) leveraging external knowledge bases (KBs), referred to as KB-VQA, emerges as a promising approach. Nevertheless, conventional unimodal retrieval techniques, which translate images into textual descriptions, often result in the loss of critical visual details. To address these challenges, this study presents two key innovations. First, we introduce fine-grained knowledge units that consist of multimodal data fragments (e.g. text fragments, entity images, and so on) in a structured manner. Rather than merely refining retrieval mechanisms, we prioritize the systematic organization and management of these knowledge units, ensuring that the structuring process itself enhances retrieval quality. Second, we propose a knowledge unit retrieval-augmented generation framework (KU-RAG) that seamlessly integrates fine-grained retrieval with MLLMs. Our KU-RAG framework not only ensures precise retrieval of relevant knowledge but also enhances reasoning capabilities through a knowledge correction chain. Experimental results demonstrate that our approach consistently outperforms existing KB-VQA methods across four benchmarks, achieving an average improvement of approximately 3% and up to 11% in the best case.

视觉问答知识检索多模态生成增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。