arXiv:2503.21401cs.ROcs.LG2025-03被引 3

让四足机器人像猫狗一样跛行,自动适应关节故障

AcL: Action Learner for Fault-Tolerant Quadruped Locomotion Control

  • 用教师-学生强化学习框架生成风格奖励,不强制模仿动作
  • 可应对单腿或双腿最多4个关节故障,实现稳定跛行
  • 适合需要高鲁棒性的真实机器人运动控制场景

四足机器人虽能学习多种运动技能,但在一个或多个关节断电时仍易失效。与猫狗受伤后可自主采用跛行步态类似,本文提出动作学习器(AcL),一种新型师生强化学习框架,使四足机器人能在多关节故障下自主调整步态以保持稳定行走。不同于传统方法强制严格模仿,AcL利用教师策略生成风格奖励,指导学生策略而不需精确复制。我们训练多个对应不同故障状态的教师策略,并通过编码器-解码器结构将其知识蒸馏为单一学生策略。相比以往仅处理单关节故障的工作,AcL支持单腿或双腿最多四个关节故障,可在故障发生时自动切换至相应跛行步态。我们在真实Go2四足机器人上验证了该方法,在单关节和双关节故障下均实现故障容错、稳定行走、平滑步态过渡及对外部扰动的鲁棒性。

原文摘要 · Abstract (English)

Quadrupedal robots can learn versatile locomotion skills but remain vulnerable when one or more joints lose power. In contrast, dogs and cats can adopt limping gaits when injured, demonstrating their remarkable ability to adapt to physical conditions. Inspired by such adaptability, this paper presents Action Learner (AcL), a novel teacher-student reinforcement learning framework that enables quadrupeds to autonomously adapt their gait for stable walking under multiple joint faults. Unlike conventional teacher-student approaches that enforce strict imitation, AcL leverages teacher policies to generate style rewards, guiding the student policy without requiring precise replication. We train multiple teacher policies, each corresponding to a different fault condition, and subsequently distill them into a single student policy with an encoder-decoder architecture. While prior works primarily address single-joint faults, AcL enables quadrupeds to walk with up to four faulty joints across one or two legs, autonomously switching between different limping gaits when faults occur. We validate AcL on a real Go2 quadruped robot under single- and double-joint faults, demonstrating fault-tolerant, stable walking, smooth gait transitions between normal and lamb gaits, and robustness against external disturbances.

四足机器人强化学习故障容错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。