arXiv:2505.20426cs.CV2025-05NeurIPS被引 7

首个系统评估大模型视角理解能力的基准,揭示其空间推理短板。

MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness

  • 构建10项任务,覆盖视角感知、推理与鲁棒性三维度
  • 43个模型测试中,多数在组合推理和扰动下表现不佳
  • 发现模型架构与链式思考提示对视角理解有关键影响

视角理解是人类视觉认知的基础,但多模态大模型(MLLMs)对视角几何的理解程度尚不明确。本文提出MMPerspective,首个系统评估MLLM视角理解能力的基准,包含2,711个真实与合成图像实例,5,083个问答对,覆盖消失点感知、透视类型推理、三维空间线条关系理解、视角保持变换下的不变性等关键能力。对43个主流MLLM的全面评估显示:模型虽在表层感知任务表现良好,但在组合推理和扰动下的空间一致性上存在明显缺陷。分析进一步揭示模型架构、规模与视角能力间的有趣关联,指出鲁棒性瓶颈及链式思考提示的增益作用。该基准为诊断与提升视觉语言系统空间理解能力提供重要工具。资源详见:https://yunlong10.github.io/MMPerspective/

原文摘要 · Abstract (English)

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understanding of perspective through 10 carefully crafted tasks across three complementary dimensions: Perspective Perception, Reasoning, and Robustness. Our benchmark comprises 2,711 real-world and synthetic image instances with 5,083 question-answer pairs that probe key capabilities, such as vanishing point perception and counting, perspective type reasoning, line relationship understanding in 3D space, invariance to perspective-preserving transformations, etc. Through a comprehensive evaluation of 43 state-of-the-art MLLMs, we uncover significant limitations: while models demonstrate competence on surface-level perceptual tasks, they struggle with compositional reasoning and maintaining spatial consistency under perturbations. Our analysis further reveals intriguing patterns between model architecture, scale, and perspective capabilities, highlighting both robustness bottlenecks and the benefits of chain-of-thought prompting. MMPerspective establishes a valuable testbed for diagnosing and advancing spatial understanding in vision-language systems. Resources available at: https://yunlong10.github.io/MMPerspective/

多模态视角理解基准测试空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。