让机器人根据视觉和语言指令精准定位可触点与空中目标点。
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
- 融合深度信息的视觉语言模型,直接生成相机坐标系下的3D点。
- 在260万样本数据集上训练,显著提升跨场景定位准确率。
- 适用于机械臂抓取、物体摆放和移动机器人导航等真实任务。
具身智能的核心在于判断在三维空间中何处行动。本文将此需求形式化为具身定位问题——基于视觉观测和语言指令预测可执行的3D点。我们定义两类目标:可触点(触摸物体表面的3D点)与空点(自由空间中的放置、导航或方向目标)。现有视觉-语言系统多依赖RGB输入,隐式重建几何结构,限制了跨场景泛化能力,尽管机器人广泛使用RGB-D传感器。为此,我们提出SpatialPoint,一个融合结构化深度信息的视觉-语言框架,可生成相机坐标系下的3D坐标。我们构建了一个包含260万样本的RGB-D数据集,涵盖可触点与空点的问答对,用于训练与评估。大量实验证明,引入深度信息显著提升具身定位性能。进一步在三类真实任务中验证:语言引导机械臂抓取、物体放置至目标位置、移动机器人导航至目标点。
原文摘要 · Abstract (English)
Embodied intelligence fundamentally requires a capability to determine where to act in 3D space. We formalize this requirement as embodied localization -- the problem of predicting executable 3D points conditioned on visual observations and language instructions. We instantiate embodied localization with two complementary target types: touchable points, surface-grounded 3D points enabling direct physical interaction, and air points, free-space 3D points specifying placement and navigation goals, directional constraints, or geometric relations. Embodied localization is inherently a problem of embodied 3D spatial reasoning -- yet most existing vision-language systems rely predominantly on RGB inputs, necessitating implicit geometric reconstruction that limits cross-scene generalization, despite the widespread adoption of RGB-D sensors in robotics. To address this gap, we propose SpatialPoint, a spatial-aware vision-language framework with careful design that integrates structured depth into a vision-language model (VLM) and generates camera-frame 3D coordinates. We construct a 2.6M-sample RGB-D dataset covering both touchable and air points QA pairs for training and evaluation. Extensive experiments demonstrate that incorporating depth into VLMs significantly improves embodied localization performance. We further validate SpatialPoint through real-robot deployment across three representative tasks: language-guided robotic arm grasping at specified locations, object placement to target destinations, and mobile robot navigation to goal positions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。