用重建场景合成视觉语言运动数据,让机器人学会真实环境中的动作。
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

- 通过3D高斯泼溅重建环境,自动生成带视觉、语言和运动轨迹的合成数据
- 在48,000条数据上训练出可预测全身动作的策略模型,实现在Unitree G1上的导航与搬物
- 无需人工标注,直接提升从仿真到现实的任务表现,适合具身智能研究者
基于感知的人形机器人行走与操作需将第一人称观测与任务指令映射至全身运动。学习该映射需要同步的视角图像、语言指令与机器人兼容的运动轨迹,但现有数据集无法在大规模下提供完整三元组。本文提出在重建的场景中合成视觉-语言-运动(VLK)监督信号。其流程利用3D高斯泼溅重建真实尺度室内环境,结合特权场景信息生成导航与物体交互轨迹,并事后渲染配对的第一人称观测。共生成48,000条无须人工干预的配对轨迹,训练一个预测短时序全身运动轨迹的VLK策略。整个身体追踪器将预测转化为物理人形机器人的动作。在真实世界单位树G1上评估导航与单物体搬运任务,结果表明:在重建场景中合成的交互能有效提供从仿真到现实的感知驱动人形机器人运动的监督信号。
原文摘要 · Abstract (English)
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker converts these predictions into actions on the physical humanoid. We evaluate on the physical Unitree G1 performing navigation and single-object transport, demonstrating that synthesized interactions in reconstructed scenes provide effective supervision for sim-to-real perception-based humanoid loco-manipulation. Project Website: https://vision-language-kinematics.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。