arXiv:2606.01458cs.RO2026-06被引 1

用逼真3D背景生成机器人操作数据,无需真人遥控就能训练出更强的类人机器人

LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World

论文配图:LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World
图 1 · 摘自论文原文
  • 用3D高斯点云构建逼真环境,自动合成带标签的操作演示数据
  • 在真实机器人上测试,自动生成数据训练的模型表现优于真人示范数据
  • 可低成本重渲染不同场景,提升模型应对环境变化的能力

类人机器人运动操作的视觉-语言-动作(VLA)策略训练受限于真人遥控示范数据采集成本高、难度大。此前基于模拟器微调的VLA策略难以有效迁移到真实任务中。本文提出LEGS(通过具身高斯点云实现运动操作),一种混合仿真环境:将机器人、物体等前景网格叠加在由手持设备捕获的场景重建的3D高斯点云(3DGS)背景之上。LEGS利用程序化运动基元生成器大规模合成带标签演示数据,无需真人遥控;并通过确定性两阶段颜色校准,使渲染图像与机器人部署摄像头对齐。在Unitree G1类人机器人上,针对三个逐渐增加全身协调难度的抓取放置任务,使用三种VLA主干网络(psi_0, pi_0.5, GR00T N1.6),仅在LEGS数据上训练的策略在所有实验中表现匹配或超越真人示范数据训练的模型。其还优于仅使用网格背景的基线,证明逼真渲染是合成数据迁移的关键。在LEGS中,机器人运动与场景外观独立记录,使得相同自动生成的演示可在新背景和物体网格下重渲染,以低于真人遥控15倍以上的成本覆盖新场景,从而增强训练数据鲁棒性。在对象与场景外观同时变化的情况下,基于重渲染LEGS-AUG数据训练的策略仍保持任务成功率,而基于真人数据训练的基线完全失败。

原文摘要 · Abstract (English)

Training vision-language-action (VLA) policies for humanoid loco-manipulation is constrained by the high cost and complexity of collecting human teleoperation demonstrations. VLA policies fine-tuned in simulators have, until now, failed to transfer effectively in humanoid loco-manipulation tasks. We present LEGS (Loco-manipulation via Embodied Gaussian Splatting), a hybrid simulator that composites a mesh foreground (robot, objects, props) over a photorealistic 3D Gaussian Splatting (3DGS) background reconstructed from a handheld scene capture. LEGS uses a procedural motion-primitive generator to synthesize labeled demonstrations at scale without human teleoperation, and a deterministic two-stage color calibration to align the rendered 3DGS image to the robot's deployment camera. On a Unitree G1 humanoid robot, across three pick-and-place tasks of increasing whole-body difficulty and three VLA backbones (psi_0, pi_0.5, GR00T N1.6), a policy trained purely on LEGS data matches or exceeds one trained on human teleoperation demos on every experiment. It also outperforms a mesh-only simulation baseline that ablates the effect of the 3DGS background, showing that photorealistic rendering is a key enabler for synthetic data transfer. Humanoid motion is recorded independently of scene appearance in LEGS, allowing the same auto-generated demonstrations to be re-rendered under new backgrounds and object meshes--covering a new scene at more than 15x lower cost than teleoperation--to augment training data for robustness to scene variations. Under combined object-and-scene appearance shift, the policy trained on re-rendered LEGS-AUG data maintains task success while the baseline trained on teleoperation data fails entirely. Our project page is located at https://legsvla.github.io/.

类人机器人仿真训练高斯点云自主生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。