用知识图谱和场景图融合提升视觉问答的准确性与细节理解
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
- 通过查询作为语义桥梁,逐步融合常识图与场景图
- 在FVQA 2.0+和MVQA上显著超越现有方法
- 适合需要高精度视觉推理的多模态应用
面向视觉问答的多模态大模型常面临双重缺陷:知识幻觉与细粒度视觉感知不足。我们发现,常识图谱和场景图恰好分别提供外部知识与细粒度视觉信息,具有互补性。然而,以往工作通常孤立处理二者,忽视其协同潜力。为此,我们提出KG-ViP框架,通过创新的检索-融合管道,以查询为语义桥梁,逐步整合两类图谱,生成统一结构化上下文,促进可靠的多模态推理。在FVQA 2.0+和MVQA基准上的大量实验表明,KG-ViP显著优于现有VQA方法。
原文摘要 · Abstract (English)
Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perception. Crucially, we identify that commonsense graphs and scene graphs provide precisely complementary solutions to these respective deficiencies by providing rich external knowledge and capturing fine-grained visual details. However, prior works typically treat them in isolation, overlooking their synergistic potential. To bridge this gap, we propose KG-ViP, a unified framework that empowers MLLMs by fusing scene graphs and commonsense graphs. The core of the KG-ViP framework is a novel retrieval-and-fusion pipeline that utilizes the query as a semantic bridge to progressively integrate both graphs, synthesizing a unified structured context that facilitates reliable multi-modal reasoning. Extensive experiments on FVQA 2.0+ and MVQA benchmarks demonstrate that KG-ViP significantly outperforms existing VQA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。