arXiv:2607.00020cs.RO2026-07

构建空间关系数据集,提升机器人视觉语言模型对物体空间布局的理解。

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories

论文配图:EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories
图 1 · 摘自论文原文
  • 用有向三元组表示物体间空间关系,支持精确评估空间结构。
  • 包含60K+操作帧和12万+相机特定场景图,覆盖多视角真实与仿真数据。
  • 实验证明现有模型在深度和视角依赖关系上仍有显著不足。

空间定位仍是视觉-语言-动作(VLA)系统在机器人操作中的关键瓶颈。尽管当前模型能识别物体并执行语言指令,却缺乏对物体空间排列的显式表征,如支撑、包含、顺序、遮挡及深度敏感关系。我们提出EmbodimentSemantic,一个用于评估具身操作中关系定位的空间场景图数据集与基准。该数据集将场景表示为有向的物体-关系-物体三元组,每条三元组使用固定关系集明确描述一对物体之间的空间关系,实现对物体绑定、关系预测与空间一致性的直接评估。数据集包含使用低成本SO101机械臂采集的真实世界操作观测,并生成对应场景图以研究实际机器人场景中的空间定位。为提供可控验证,我们还引入基于模拟器的LIBERO基准,涵盖超过60,000个操作帧和超过120,000个相机特定场景图,分别来自第三人称与腕部视角,其真实关系由MuJoCo几何、世界坐标、相机投影及可见性约束自动生成。我们进一步测试将场景图注入现有视觉语言策略提示是否可提升下游控制性能。在开源与商用视觉语言模型上的实验表明,当前模型虽常能预测合理关系,但在精确的深度感知与视角依赖空间结构上仍表现不佳。EmbodimentSemantic为诊断视觉语言模型感知中的空间定位问题提供了统一框架,并检验其在具身操作中的应用价值。

原文摘要 · Abstract (English)

Spatial grounding remains a key limitation of vision-language-action (VLA) systems for robotic manipulation. While current models can recognize objects and follow language instructions, they often lack an explicit representation of how objects are arranged in space, including support, containment, ordering, occlusion, and depth-sensitive relations. We introduce EmbodimentSemantic, a spatial scene-graph dataset and benchmark for evaluating relational grounding in embodied manipulation. EmbodimentSemantic represents scenes as directed object-relation-object triplets, where each triplet specifies a spatial relation between an ordered pair of objects using a fixed set of relations. This representation enables direct evaluation of object binding, relation prediction, and spatial consistency. The dataset includes real-world manipulation observations collected with the low-cost SO101 robot arm, together with generated scene graphs for studying spatial grounding in practical robotic settings. To provide controlled validation, we also introduce a simulator-grounded LIBERO benchmark with over 60K manipulation frames and more than 120K camera-specific scene graphs across paired third-person and wrist views, where ground-truth relations are derived automatically from MuJoCo geometry, world coordinates, camera projections, and visibility constraints. We further test whether scene graphs improve downstream control by injecting them into existing VLA policy prompts. Experiments across open-source and commercial VLMs show that current models often predict plausible relations but struggle with exact depth-aware and viewpoint-dependent spatial structure. EmbodimentSemantic provides a unified framework for diagnosing spatial grounding in VLM perception and testing its utility for VLA manipulation.

空间理解机器人操作视觉语言模型场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。