用结构化场景图增强视觉问答,提升复杂图像中物体识别精度
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
- 引入结构化场景图辅助多模态大模型理解图像
- 在VG-150和AUG数据集上显著提升物体识别与定位准确率
- 适合需要精确空间理解的视觉分析任务
多模态大语言模型(MLLMs)如GPT-4o、Gemini、LLaVA和Flamingo在整合视觉与文本模态方面取得显著进展,擅长视觉问答(VQA)、图像描述生成和内容检索等任务。然而,在复杂场景中,面对重叠或细小物体时,仍难以准确识别、计数并确定其空间位置。为此,本文提出一种基于多模态检索增强生成(RAG)的新框架,引入结构化场景图以增强对图像中物体识别、关系判断和空间理解的能力。该框架显著提升了模型在需精准视觉描述任务中的表现,尤其在俯视视角或物体密集排列的场景下。我们在专注于第一人称视觉理解的VG-150数据集和涉及航空影像的AUG数据集上进行了大量实验,结果表明,该方法在各类VQA任务中持续优于现有MLLMs,展现出更强的物体识别、定位与量化能力,并提供更准确的视觉描述。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning, and content retrieval. They can generate coherent and contextually relevant descriptions of images. However, they still face challenges in accurately identifying and counting objects and determining their spatial locations, particularly in complex scenes with overlapping or small objects. To address these limitations, we propose a novel framework based on multimodal retrieval-augmented generation (RAG), which introduces structured scene graphs to enhance object recognition, relationship identification, and spatial understanding within images. Our framework improves the MLLM's capacity to handle tasks requiring precise visual descriptions, especially in scenarios with challenging perspectives, such as aerial views or scenes with dense object arrangements. Finally, we conduct extensive experiments on the VG-150 dataset that focuses on first-person visual understanding and the AUG dataset that involves aerial imagery. The results show that our approach consistently outperforms existing MLLMs in VQA tasks, which stands out in recognizing, localizing, and quantifying objects in different spatial contexts and provides more accurate visual descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。