arXiv:2509.12858cs.RO2025-09被引 4

让纯本体感知的机器人具备预见性,无需依赖脆弱的视觉系统。

Contrastive Representation Learning for Robust Sim-to-Real Transfer of Adaptive Humanoid Locomotion

  • 用对比学习让机器人从仿真中‘学’到环境信息,注入本体感知策略。
  • 实现在30厘米高台阶和26.5度斜坡上零样本迁移的稳定行走。
  • 适合追求鲁棒性与自适应性的真实人形机器人部署者。

强化学习在人形机器人行走方面取得了显著进展,但现实部署仍面临根本矛盾:策略需在反应式本体感知控制的鲁棒性与复杂易损的感知驱动系统的主动性之间二选一。本文提出一种新范式,使纯本体感知策略具备前瞻性能力,实现感知的预见性而不承担其运行时成本。核心贡献是一个对比学习框架,迫使智能体的隐状态编码来自仿真的特权环境信息。关键的是,这种‘提炼出的觉知’赋能自适应步态时钟,使策略能基于对地形的推断主动调整步伐节奏。该协同机制破解了传统刚性时钟步态与不稳定的无时钟策略之间的权衡。我们在全尺寸人形机器人上验证了零样本的仿真到现实迁移,成功实现复杂地形上的高度鲁棒行走,包括30厘米高的台阶和26.5°的斜坡,证明了方法的有效性。

原文摘要 · Abstract (English)

Reinforcement learning has produced remarkable advances in humanoid locomotion, yet a fundamental dilemma persists for real-world deployment: policies must choose between the robustness of reactive proprioceptive control or the proactivity of complex, fragile perception-driven systems. This paper resolves this dilemma by introducing a paradigm that imbues a purely proprioceptive policy with proactive capabilities, achieving the foresight of perception without its deployment-time costs. Our core contribution is a contrastive learning framework that compels the actor's latent state to encode privileged environmental information from simulation. Crucially, this ``distilled awareness" empowers an adaptive gait clock, allowing the policy to proactively adjust its rhythm based on an inferred understanding of the terrain. This synergy resolves the classic trade-off between rigid, clocked gaits and unstable clock-free policies. We validate our approach with zero-shot sim-to-real transfer to a full-sized humanoid, demonstrating highly robust locomotion over challenging terrains, including 30 cm high steps and 26.5° slopes, proving the effectiveness of our method. Website: https://lu-yidan.github.io/cra-loco.

人形机器人强化学习仿真实现自适应步态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。