arXiv:2411.16537cs.CVcs.AI2024-11CVPR被引 185

构建机器人空间理解数据集,提升视觉语言模型的场景感知能力

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

论文配图:RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
图 1 · 摘自论文原文
  • 构建包含100万张图像与5000个3D扫描的多模态数据集
  • 在空间关系预测等任务上显著优于基线模型
  • 适合研究机器人视觉、空间推理与多模态学习的研究者

空间理解是机器人感知环境、推理并有意义交互的关键能力。当前机器人依赖视觉语言模型实现该能力,但这些模型在空间推理任务中表现受限,因其训练数据来自缺乏复杂空间理解的通用图像数据集,例如未充分涵盖参考系认知(如以自我、世界或物体为中心的视角)。为此,我们提出RoboSpatial,一个大规模机器人空间理解数据集,包含真实室内与桌面场景的3D扫描和头戴式图像,标注了300万条与机器人任务相关的空间关系。数据集融合2D头戴图像与3D扫描,支持2D与3D双重使用。实验表明,基于RoboSpatial训练的模型在空间可操作性预测、空间关系预测及机器人操作任务中均优于基线模型。

原文摘要 · Abstract (English)

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant challenges in spatial reasoning tasks, as their training data are based on general-purpose image datasets that often lack sophisticated spatial understanding. For example, datasets frequently do not capture reference frame comprehension, yet effective spatial reasoning requires understanding whether to reason from ego-, world-, or object-centric perspectives. To address this issue, we introduce RoboSpatial, a large-scale dataset for spatial understanding in robotics. It consists of real indoor and tabletop scenes, captured as 3D scans and egocentric images, and annotated with rich spatial information relevant to robotics. The dataset includes 1M images, 5k 3D scans, and 3M annotated spatial relationships, and the pairing of 2D egocentric images with 3D scans makes it both 2D- and 3D- ready. Our experiments show that models trained with RoboSpatial outperform baselines on downstream tasks such as spatial affordance prediction, spatial relationship prediction, and robot manipulation.

空间理解视觉语言模型机器人多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。