新基准M$^3$-VQA挑战多实体多跳视觉问答,推动模型精细理解与复杂推理。
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

- 设计多实体多跳问题,融合图文源信息进行复杂推理
- 模型在无外部知识时表现差,有精确证据后显著提升
- 推理驱动的检索优于传统方法,适合研究多模态推理者
我们提出M$^3$-VQA,一个面向多模态大语言模型(MLLMs)的新型知识增强型视觉问答基准,旨在提升对细粒度多模态实体理解和复杂多跳推理的评估能力。与现有聚焦粗粒度类别和单实体简单推理的VQA数据集不同,M$^3$-VQA引入了来自视觉与文本源的多种多实体问题,要求模型在多文档间执行串行与并行多跳推理,并基于可追溯、详细的证据及精心构建的多模态知识库。我们在三种设置下评估16个领先MLLMs:无外部知识、使用黄金证据、以及检索增强输入。结果表明,模型在缺乏外部信息时表现不佳,但获得精确证据后显著改善。此外,推理感知的代理式检索优于启发式方法,凸显结构化推理对复杂多模态理解的重要性。M$^3$-VQA为推进MLLMs的多模态推理能力提供了更具挑战性的评估框架。代码与数据集已开源:https://github.com/CASIA-IVA-Lab/M3VQA。
原文摘要 · Abstract (English)
We present M$^3$-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multimodal entity understanding and complex multi-hop reasoning. Unlike existing VQA datasets that focus on coarse-grained categories and simple reasoning over single entities, M$^3$-VQA introduces diverse multi-entity questions involving multiple distinct entities from both visual and textual sources. It requires models to perform both sequential and parallel multi-hop reasoning across multiple documents, supported by traceable, detailed evidence and a curated multimodal knowledge base. We evaluate 16 leading MLLMs under three settings: without external knowledge, with gold evidence, and with retrieval-augmented input. The poor results reveal significant challenges for MLLMs in knowledge acquisition and reasoning. Models perform poorly without external information but improve markedly when provided with precise evidence. Furthermore, reasoning-aware agentic retrieval surpasses heuristic methods, highlighting the importance of structured reasoning for complex multimodal understanding. M$^3$-VQA presents a more challenging evaluation for advancing the multimodal reasoning capabilities of MLLMs. Our code and dataset are available at https://github.com/CASIA-IVA-Lab/M3VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。