通过融合细粒度视觉与语言特征,提升复杂视觉问答的推理能力。
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
- 结合全局特征与物体检测、场景图等细粒度视觉信息进行多模态融合
- 在GQA上达到77.5%准确率,优于现有大模型基线
- 适合需要深度视觉与概念推理的研究者和应用开发
复杂视觉问答(Complex VQA)任务要求复杂的多模态推理与外部知识整合,现有大视觉语言模型(LVLMs)常因依赖高层全局特征而受限。为此,我们提出MV-CoRe(多模态视觉-概念推理)模型,通过深度融合预训练视觉大模型(VLMs)与语言大模型(LLMs)的全局嵌入,以及物体检测特征和场景图表示等细粒度语义感知视觉特征,显著提升复杂推理能力。创新的多模态融合变压器(Multimodal Fusion Transformer)对多种特征进行深度交互,实现丰富的跨模态注意力。我们在GQA、A-OKVQA和OKVQA等挑战性基准上评估了该模型,训练基于VQAv2。实验表明,MV-CoRe在各项任务中持续超越主流LVLM基线,在GQA上整体准确率达77.5%。消融实验验证了物体特征与场景图特征的关键作用,人工评估也进一步证明其在事实正确性和推理深度上的优势,展现出强大的深层视觉与概念理解能力。
原文摘要 · Abstract (English)
Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for existing large vision-language models (LVLMs) often limited by their reliance on high-level global features. To address this, we propose MV-CoRe (Multimodal Visual-Conceptual Reasoning), a novel model designed to enhance Complex VQA performance through the deep fusion of diverse visual and linguistic information. MV-CoRe meticulously integrates global embeddings from pre-trained Vision Large Models (VLMs) and Language Large Models (LLMs) with fine-grained semantic-aware visual features, including object detection characteristics and scene graph representations. An innovative Multimodal Fusion Transformer then processes and deeply integrates these diverse feature sets, enabling rich cross-modal attention and facilitating complex reasoning. We evaluate MV-CoRe on challenging Complex VQA benchmarks, including GQA, A-OKVQA, and OKVQA, after training on VQAv2. Our experimental results demonstrate that MV-CoRe consistently outperforms established LVLM baselines, achieving an overall accuracy of 77.5% on GQA. Ablation studies confirm the critical contribution of both object and scene graph features, and human evaluations further validate MV-CoRe's superior factual correctness and reasoning depth, underscoring its robust capabilities for deep visual and conceptual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。