构建新数据集揭示大模型空间推理能力短板
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
- 用边界框和场景图增强空间推理能力
- 人类视角问题比相机视角更难回答
- 思维链提示对复杂空间推理无效
大型多模态模型在视觉语言任务中表现强劲,但其空间推理能力尚未充分研究。本文构建了新的视觉问答数据集Spatial-MM,系统评估大模型的空间理解与推理能力。分析显示:边界框和场景图(包括合成数据)能显著提升空间推理表现;模型在人类视角问题上的表现明显弱于相机视角;思维链提示对涉及空间关系的多跳问题无改善作用;多模态大模型在空间推理步骤上的准确率远低于非空间步骤。对GQA-spatial的扰动分析表明,模型在基础物体检测上强于复杂空间推理。本研究为后续空间推理研究提供了基准数据集与深入洞见。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities. Our analyses on object-relationship and multi-hop reasoning reveal several important findings. Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning. Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image. Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations. % Moreover, spatial reasoning steps are much less accurate than non-spatial ones across MLLMs. Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning. We believe our benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning. Spatial-MM benchmark is available at: https://github.com/FatemehShiri/Spatial-MM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。