首个3D空间推理综合基准,评估模型对三维场景的理解能力。
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

- 构建12类问题、2772个标注的3D空间推理数据集
- 发现大模型在罕见视角下性能显著下降,多对象推理存在缺陷
- 适合研究三维感知、机器人、AR/VR领域模型的开发者
3D空间推理是指分析和理解三维空间中物体的位置、朝向及空间关系的能力,有助于模型全面理解三维场景,广泛应用于自动驾驶、机器人和增强现实/虚拟现实等领域。尽管大规模多模态模型在图像和视频理解任务中取得显著进展,但其在多样自然图像上的3D空间推理能力仍缺乏系统研究。本文提出首个综合性3D空间推理基准3DSRBench,包含2,772个手动标注的视觉问答对,覆盖12种问题类型。通过平衡数据分布并采用新颖的FlipEval评估策略,对多种开源与专有大模型进行了全面评测。结果揭示了模型在高度、朝向、位置及多对象推理方面的局限性,尤其在非常见6D视角图像上表现退化。该基准为未来具备强空间推理能力的大模型发展提供了重要洞察。
原文摘要 · Abstract (English)
3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。