arXiv:2602.06382cs.RO2026-02被引 9

从原始像素端到端训练人形机器人,克服视觉噪声与地形差异挑战

Now You See That: Learning End-to-End Humanoid Locomotion from Raw Pixels

  • 用高保真深度仿真模拟真实传感器误差,提升训练真实性
  • 通过隐空间对齐和抗噪辅助任务,实现从高清地图到噪声深度图的知识迁移
  • 结合多评价值函数与多判别器,让机器人自适应不同地面类型

由于模拟到现实的差距引入显著感知噪声,且在不同地形上训练统一策略受冲突学习目标阻碍,视觉驱动的人形机器人行走仍具挑战。本文提出端到端视觉驱动行走框架。为增强模拟到现实的鲁棒性,开发了包含立体匹配伪影和校准不确定性的高保真深度传感器仿真;进一步提出视觉感知行为蒸馏方法,通过隐空间对齐与抗噪辅助任务,实现从理想高度图向噪声深度观测的有效知识迁移。针对多样地形适应,引入地形特异性奖励塑造,结合多评价值函数与多判别器学习,使专用网络捕捉各地形的动力学特性与运动先验。在搭载不同立体深度相机的两款人形机器人平台上验证,所提策略在复杂环境(如高平台、宽缝隙)及精细任务(如双向长阶梯行走)中表现稳健。

原文摘要 · Abstract (English)

Achieving robust vision-based humanoid locomotion remains challenging due to two fundamental issues: the sim-to-real gap introduces significant perception noise that degrades performance on fine-grained tasks, and training a unified policy across diverse terrains is hindered by conflicting learning objectives. To address these challenges, we present an end-to-end framework for vision-driven humanoid locomotion. For robust sim-to-real transfer, we develop a high-fidelity depth sensor simulation that captures stereo matching artifacts and calibration uncertainties inherent in real-world sensing. We further propose a vision-aware behavior distillation approach that combines latent space alignment with noise-invariant auxiliary tasks, enabling effective knowledge transfer from privileged height maps to noisy depth observations. For versatile terrain adaptation, we introduce terrain-specific reward shaping integrated with multi-critic and multi-discriminator learning, where dedicated networks capture the distinct dynamics and motion priors of each terrain type. We validate our approach on two humanoid platforms equipped with different stereo depth cameras. The resulting policy demonstrates robust performance across diverse environments, seamlessly handling extreme challenges such as high platforms and wide gaps, as well as fine-grained tasks including bidirectional long-term staircase traversal.

人形机器人视觉导航强化学习仿真实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。