用合成世界训练机器人看懂空间距离,为人机交互打基础
Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds
- 在Omniverse中生成带真实位姿的合成数据,支持视觉语言模型学习空间推理
- 通过4×4变换矩阵标注物体姿态,重点训练对Z轴距离的感知能力
- 数据集开源,适合研究具身智能与人机交互的学者使用
我们提出一个概念框架,用于训练视觉语言模型(VLM)完成视觉视角转换(VPT),这是实现具身认知的核心能力,对人机交互至关重要。作为迈向该目标的第一步,我们引入了一个在NVIDIA Omniverse中生成的合成数据集,支持空间推理任务的监督学习。每个样本包含一幅RGB图像、一段自然语言描述以及一个表示物体姿态的真实4×4变换矩阵。本研究聚焦于推断Z轴距离这一基础技能,未来可扩展至完整的6自由度(DOFs)空间推理。该数据集已公开,以支持后续研究。此项工作为实现具备空间理解能力的交互式机器人系统奠定了基础。
原文摘要 · Abstract (English)
We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal, we introduce a synthetic dataset, generated in NVIDIA Omniverse, that enables supervised learning for spatial reasoning tasks. Each instance includes an RGB image, a natural language description, and a ground-truth 4X4 transformation matrix representing object pose. We focus on inferring Z-axis distance as a foundational skill, with future extensions targeting full 6 Degrees Of Freedom (DOFs) reasoning. The dataset is publicly available to support further research. This work serves as a foundational step toward embodied AI systems capable of spatial understanding in interactive human-robot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。