测试视觉语言模型跨视角整合能力,发现其3D空间理解普遍薄弱。
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

- 设计新基准,要求模型从多视角构建全局3D场景认知
- 主流模型在3D空间关系上表现差,仅2D平面关系良好
- 适合研究多视角融合、机器人感知等场景的开发者参考
当前视觉语言模型(VLM)评测大多聚焦单视图或有限视图感知,未检验将多视角观察整合为连贯世界中心(非相对)3D心智模型的核心认知能力。我们提出MultiView-Bench,一个专门用于诊断多视角整合能力的基准,要求模型摆脱瞬时视角干扰,在固定全局坐标系中定位物体,这是机械零件装配等下游任务的前提。对前沿VLM系统的系统评估显示:模型在单图像2D平面关系上表现良好,但在3D空间关系和跨视角信息聚合上存在显著困难。此外还发现模型对非标准轴向敏感,且受物体颜色与纹理变化影响。针对此,我们提出ViewNavigator,通过主动视角选择与证据融合,使四个基础模型在六图像限制下提升12.3–20.0个百分点;预算扩展后,GPT-5最高提升达27个百分点。
原文摘要 · Abstract (English)
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3--20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。