构建首个6维空间推理基准,测试大模型在三维位置与朝向上的理解能力。
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
- 设计包含4类能力的合成数据集,覆盖多物体识别到6维空间推理
- 7种题型5级难度下,模型性能随复杂度提升显著下降,3D推理最弱
- 提出相对性能衰减率,揭示模型在真实图像中也存在类似偏差
尽管大型多模态模型(LMMs)在视觉场景解析与推理方面表现出色,但其对复杂且精确的三维空间推理能力仍不明确。现有基准主要关注二维空间理解,缺乏全面评估六维空间推理的框架。为此,我们提出Spatial457,一个可扩展、无偏见的合成数据集,具备四大空间推理能力:多物体识别、二维位置、三维位置和三维方向。我们构建了分层评估结构,涵盖7种题型、5个难度等级,从基础单物体识别延伸至新提出的复杂6维空间推理任务。在PulseCheck457上评估多种LMMs,发现性能随任务复杂度上升普遍下降,尤其在3D推理与6维空间任务中表现不佳。为量化此挑战,我们引入相对性能衰减率(RPDR),凸显3D推理中的关键短板。借助数据集的无偏设计,我们还发现不同属性间的预测偏差,其模式在真实图像设置中亦可复现。代码与数据已开源于https://github.com/XingruiWang/Spatial457。
原文摘要 · Abstract (English)
Although large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks focus predominantly on 2D spatial understanding and lack a framework to comprehensively evaluate 6D spatial reasoning across varying complexities. To address this limitation, we present Spatial457, a scalable and unbiased synthetic dataset designed with 4 key capability for spatial reasoning: multi-object recognition, 2D location, 3D location, and 3D orientation. We develop a cascading evaluation structure, constructing 7 question types across 5 difficulty levels that range from basic single object recognition to our new proposed complex 6D spatial reasoning tasks. We evaluated various large multimodal models (LMMs) on PulseCheck457, observing a general decline in performance as task complexity increases, particularly in 3D reasoning and 6D spatial tasks. To quantify these challenges, we introduce the Relative Performance Dropping Rate (RPDR), highlighting key weaknesses in 3D reasoning capabilities. Leveraging the unbiased attribute design of our dataset, we also uncover prediction biases across different attributes, with similar patterns observed in real-world image settings. The code and data are released in https://github.com/XingruiWang/Spatial457.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。