arXiv:2510.18348cs.ROcs.AI2025-10中稿 · IROS 2026

让机器人通过感知地形自适应行走,无需预设步态模式。

PGTT: Phase-Guided Terrain Traversal for Perceptive Legged Locomotion

  • 用相位引导的奖励机制替代固定步态约束,提升适应性。
  • 在楼梯和障碍物上成功率比最优基线高7.5%和9%。
  • 适合需跨平台部署的足式机器人,免去重新调参。

当前先进的感知强化学习控制器要么依赖振荡器或逆运动学步态先验,限制动作空间、引入偏差且难以适配不同机器人结构;要么完全依赖视觉反馈,无法提前预判后腿地形,对观测噪声敏感。本文提出相位引导的地形穿越方法(PGTT),通过奖励设计实现步态结构引导,降低对先验动作空间的依赖。PGTT将每条腿的相位编码为三次埃尔米特样条,根据局部高度图统计信息动态调整抬腿高度,并在摆动阶段引入接触惩罚。策略直接在关节空间中执行,支持形态无关部署。在MuJoCo(MJX)中,基于程序化生成的台阶状地形,结合课程学习与域随机化训练,PGTT在受推力扰动下成功率中位数比次优基线高出7.5%,在离散障碍物上高出9%,同时保持相近的速度追踪性能。在Unitree Go2上通过实时激光雷达高程转高度图管道验证,初步结果表明相同超参数下可在ANYmal-C上复现,证明了地形自适应、相位引导的奖励设计具有跨平台迁移能力,无需特定平台的策略先验或大量重调参。

原文摘要 · Abstract (English)

State-of-the-art perceptive Reinforcement Learning controllers for legged robots typically either (i) impose oscillator-or IK-based gait priors that constrain the action space, bias policy optimization, and limit adaptability across robot morphologies, or (ii) operate "blind," making them unable to anticipate hind-leg terrain and brittle to observation noise. We propose Phase-Guided Terrain Traversal (PGTT), a perception-aware deep-RL approach that enforces gait structure through reward shaping, thereby reducing inductive bias compared to oscillator- or IK-conditioned action priors. PGTT encodes per-leg phase as a cubic Hermite spline, adapts swing height to local heightmap statistics, and adds a swing-phase contact penalty, while the policy acts directly in joint space for morphology-agnostic deployment. Trained in MuJoCo (MJX) on procedurally generated stair-like terrains with curriculum learning and domain randomization, PGTT achieves the highest success rate among the evaluated baselines under push disturbances (median +7.5% over the next-best baseline) and on discrete obstacles (+9%), while maintaining comparable velocity tracking. We validate PGTT on a Unitree Go2 using a real-time LiDAR elevation-to-heightmap pipeline and report preliminary results on ANYmal-C using the same hyperparameters. These results provide early evidence that terrain-adaptive, phase-guided reward shaping can transfer across platforms without platform-specific policy priors or extensive re-tuning.

足式机器人强化学习地形适应奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。