通过分析步态相位中的有效秩,提升足式机器人仿真到现实的平滑性。
Mind the Phase: Effective Rank and Representation Health in Legged Locomotion

- 按步态相位分析策略雅可比矩阵的有效秩,揭示隐藏的表征结构。
- 标准架构在摆动期比支撑期多出约两个有效秩维度,而普通MLP无此差异。
- 基于此设计训练优化方案,使真实机器人关节抖动降低约3倍。
强化学习已成为足式机器人运动控制的主流范式,通过大规模并行仿真实现后空翻、障碍穿越等复杂行为。在PPO算法非平稳性影响下,浅层网络仍是默认架构,依赖精心设计的课程和环境,但策略所学表征仍不清晰,缺乏训练阶段对硬件表现的预测信号。本文通过分析策略雅可比矩阵的有效秩,发现将秩条件化于步态相位可揭示全局平均掩盖的架构结构。具体而言,标准架构(如层归一化与残差连接)在摆动相比支撑相多分配约两个有效秩维度,而纯MLP则无此差异。基于此,我们提出一种简单方法,将这些表征特征转化为更平滑、可靠的仿真到现实迁移。实际应用中,该方法使物理Spot机器人关节抖动降低约3倍,表明表征健康度是追踪仿真到现实平滑性的有效训练期指标。
原文摘要 · Abstract (English)
Reinforcement learning has become the leading paradigm in legged locomotion, enabling complex behaviors from backflips to parkour through massively parallel simulation. Under PPO's non-stationarity, shallow networks remain the de facto architecture, supported by carefully staged curricula and environments, yet the representations these policies learn stay poorly understood, leaving no training-time signal of how they will behave on hardware. In this work, we empirically study locomotion policies through the effective rank of the policy Jacobian and show that conditioning rank on the gait phase exposes architectural structure that global rank averages away. In particular, we find that standard architectural choices, namely layer normalization and residual connections, allocate roughly two more dimensions of effective rank to swing than to stance, which is fully absent in vanilla MLPs. Building on this, we propose a simple recipe that turns these representational signatures into smoother, more reliable sim-to-real transfer. In practice, this results in roughly 3x lower joint jitter that holds from simulation onto a physical Spot, suggesting that representation health is an effective training-time lens to track sim-to-real smoothness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。