评测视觉语言模型在3D环境中的基础空间感知能力,发现其交互感知仍脆弱。
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

- 构建6类任务的机器人导向3D基准,涵盖空间理解与交互感知
- 13个模型测试显示高阶空间推理强但交互感知差,需增强3D先验
- 推出130万条数据集,微调后显著提升低层空间智能表现
当前视觉语言模型是否具备在复杂3D环境中理解并推理具身交互的能力?我们提出Embodied3DBench,一个面向具身3D环境低层空间智能的机器人中心型基准。为系统评估这些基础感知能力,该基准包含6个任务类别,分为两大核心组:空间结构理解(定位、空间关系预测、多视角对应)和交互感知(可操作性预测、抓取点预测、轨迹预测)。基准覆盖12个子类别,包含超过21,000个高质量问答对。我们评估了13个前沿模型,结果表明:尽管当前模型在高层次空间推理(如物体间位置关系)上表现良好,但在交互感知方面仍显脆弱,暴露出缺乏稳健的3D交互先验。为弥补这一能力差距,我们进一步构建了一个包含130万条问答对的大规模训练数据集。值得注意的是,在该数据集上微调能显著提升低层空间智能。最终,Embodied3DBench不仅提供系统化评估框架,还提供可扩展的数据解决方案,为发展交互感知多模态系统设定了明确目标。
原文摘要 · Abstract (English)
Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelligence in embodied 3D environments. To systematically evaluate these foundational perceptual capabilities, the benchmark includes 6 task categories divided into two core groups: Spatial Structural Understanding (Grounding, Spatial Relation Prediction, and Multi-view Correspondence) and Interaction-Oriented Perception (Affordance Prediction, Grasp Point Prediction, and Trajectory Prediction). The benchmark spans 12 subcategories and contains over 21k high-quality question-answer pairs. We evaluate 13 state-of-the-art models, and the results show that while current models exhibit relatively strong high-level spatial reasoning, such as understanding object-to-object positional relations, they remain fragile in interaction-oriented perception, highlighting a significant lack of robust 3D-aware interaction priors. To actively bridge this capability gap revealed by our benchmark, we further synthesize a large-scale training dataset comprising 1.3M QA pairs. Notably, fine-tuning on this dataset yields significant improvements in low-level spatial intelligence. Ultimately, Embodied3DBench fills a critical gap by providing both a systematic evaluation framework and a scalable data solution, setting a clear target for the development of interaction-aware multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。