arXiv:2608.07328cs.ROcs.LG2026-08中稿 · the 2026 IEEE/RSJ …

让68公斤四足机器人在电机失效时自适应调整步伐,保持稳定行走。

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

论文配图:Learning Fault-Tolerant Locomotion with Adaptive Gait Timing
图 1 · 摘自论文原文
  • 用深度强化学习让机器人从本体感知中重建状态,自主应对故障。
  • 在不预设故障策略下,实现对地形和电机退化的自适应步频调节。
  • 在68公斤真机上验证,复杂地形与平坦地面均有效。

硬件故障要求腿式机器人快速重新组织协调与步态时序以维持稳定性和移动性。这对大型四足机器人尤为困难,其更大的质量与更紧的驱动限制降低了激进、高频补偿策略的可行性,而这类策略常用于小型平台。本文提出一种基于深度强化学习的电机功率损失下的容错运动方法。该方法采用非对称的演员-评论家架构,训练时评论家可访问特权信息,而演员则从本体感知观测中学习重建相应隐表示。引入隐空间对齐损失,促进演员与评论家表示的一致性。此外,通过在动作空间中加入可学习的步态频率参数,实现对地形变化和执行器退化的自适应步频调节,无需预设故障腿策略。方法在高保真仿真中于不规则地形上进行验证,并在真实世界中使用一台68公斤四足机器人于平坦地面完成实验。

原文摘要 · Abstract (English)

Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.

四足机器人容错控制强化学习步态自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。