arXiv:2608.02653cs.RO2026-08被引 1

一个策略搞定人形机器人复杂地形全身体感运动,无需预设动作或标签。

Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation

论文配图:Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation
图 1 · 摘自论文原文
  • 用强化学习扩展速度追踪策略,融合障碍交互动作实现全身体感控制。
  • 仅需少量初始动作,通过动态可行参考生成多障碍场景下的技能。
  • 自主决定何时何地使用技能,无标签、无状态机,零样本迁移至真实硬件。

现有类人机器人全身控制仍难以模拟人类在复杂地形中的灵活移动:要么依赖表达性全身参考但无法泛化,要么在线反应地形却忽略四肢、躯干和膝盖。本文提出 exttt{Light-Loco-Parkour}(LLP),一种端到端的感知式全身运动系统,仅用机载深度信息与速度指令即可部署单一策略,自主决策行走、平衡、攀爬、跨步或跳跃,并在完成任务后恢复行进。相比此前系统,有三项贡献:第一,构建了从强化学习训练的速度追踪策略扩展出包含攀爬技能的全身感知控制流程;第二,通过稀疏种子动作扩展生成适应不同障碍几何的动态可行、地形匹配参考,不依赖大规模动作库;第三,通过奖励信号学习自主技能切换机制,仅凭深度与指令判断何时及如何调用技能,无需技能标签、手写状态机或运行时动作生成器。仿真与真实实验表明,在基准地形和未见障碍变化中均具高成功率,同一策略可零样本迁移至室内外真实硬件。结果展示了仅用机载传感与单一策略实现类人机器人户外自主感知式全身运动。

原文摘要 · Abstract (English)

Existing humanoid whole-body control systems still fall short of the way humans move through cluttered terrain: they either track expressive whole-body references without terrain generalization, or react to terrain online while leaving the arms, torso, and knees largely unused. We present \texttt{Light-Loco-Parkour} (LLP), an end-to-end perceptive whole-body locomotion system that closes this gap with a single deployable policy. Conditioned only on onboard depth and a velocity command, the policy decides when to walk, balance, climb, step down, or vault, with no reference input, skill label, hand-coded gate, or runtime motion graph. Compared with prior humanoid systems, LLP makes three contributions. First, it introduces a whole-body perceptive-control pipeline that extends an RL-trained, velocity-tracking locomotion policy with parkour skills learned from object-interacting motions, so the same policy tracks velocity in open terrain, executes whole-body traversal at obstacles, and resumes locomotion afterward. Second, it acquires terrain-conditioned skills from sparse seeds by expanding a single motion into dynamically feasible, terrain-paired references across obstacle geometry, rather than relying on a large motion corpus. Third, it learns autonomous skill transitions from reward, letting the policy decide when and which whole-body skill to invoke from depth and command alone, with no one-hot skill label, hand-coded state machine, or runtime motion generator. Simulation and real-world experiments show high success across both benchmarked terrains and unseen obstacle variations, and the same policy transfers zero-shot to indoor and outdoor hardware experiments. These results demonstrate autonomous perceptive whole-body locomotion on a humanoid in outdoor settings, using only onboard sensing and a single deployable policy.

人形机器人全身运动自主决策强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。