纯视觉控制的类人机器人零样本开门,性能超人类操作员31.7%。
Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
- 分阶段重置探索+GRPO微调,提升仿真到现实的迁移稳定性。
- 仅用仿真数据训练,零样本通过多种门型,任务完成时间快31.7%。
- 首个基于纯RGB图像的类人机器人全身协同开锁控制策略。
GPU加速的逼真仿真为机器人学习提供了可扩展的数据生成路径,通过大规模物理与视觉随机化,使策略能泛化至未预设环境。基于此,我们提出一种教师-学生-自举学习框架,用于基于视觉的类人机器人运动-操作任务,以关节物体交互为高难度基准。方法引入分阶段重置探索策略,稳定长时程特权策略训练;采用基于GRPO的微调流程,缓解部分可观测性问题,提升闭环一致性。整个策略完全在仿真数据上训练,实现跨多种门型的鲁棒零样本表现,在相同全身控制架构下,任务完成时间比人类远程操作员快31.7%。这是首个仅使用纯RGB感知实现多样化关节式运动-操作的类人机器人仿真到现实迁移策略。
原文摘要 · Abstract (English)
Recent progress in GPU-accelerated, photorealistic simulation has opened a scalable data-generation path for robot learning, where massive physics and visual randomization allow policies to generalize beyond curated environments. Building on these advances, we develop a teacher-student-bootstrap learning framework for vision-based humanoid loco-manipulation, using articulated-object interaction as a representative high-difficulty benchmark. Our approach introduces a staged-reset exploration strategy that stabilizes long-horizon privileged-policy training, and a GRPO-based fine-tuning procedure that mitigates partial observability and improves closed-loop consistency in sim-to-real RL. Trained entirely on simulation data, the resulting policy achieves robust zero-shot performance across diverse door types and outperforms human teleoperators by up to 31.7% in task completion time under the same whole-body control stack. This represents the first humanoid sim-to-real policy capable of diverse articulated loco-manipulation using pure RGB perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。