arXiv:2511.06371cs.RO2025-11AAAI被引 6

让仿人机器人学会自适应切换行走、跑步等动作,应对复杂地形。

Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-Tuning

  • 先训练多种动作策略,通过多行为蒸馏融合成通用控制器。
  • 在线收集真实场景反馈,强化模型在不同地形的适应能力。
  • 在仿真和真实机器人上验证,性能优于传统独立训练方法。

仿人机器人有望学习多种类人运动行为,如站立、行走、奔跑和跳跃。然而,现有方法通常为每种技能单独训练策略,导致行为特定的控制器在不规则地形和多样化场景下泛化能力差、表现脆弱。为此,我们提出自适应仿人控制(AHC),采用两阶段框架,在不同技能和地形间学习自适应运动控制。首先,训练多个基础运动策略,并通过多行为蒸馏获得一个基础多行为控制器,实现基于环境的自适应行为切换;随后,在更多样地形上执行自适应行为并收集在线反馈,进行强化微调,提升控制器的地形适应性。我们在仿真和真实世界中对Unitree G1机器人进行了实验。结果表明,该方法在各种情境和地形下均表现出强适应性。项目网站:https://ahc-humanoid.github.io。

原文摘要 · Abstract (English)

Humanoid robots are promising to learn a diverse set of human-like locomotion behaviors, including standing up, walking, running, and jumping. However, existing methods predominantly require training independent policies for each skill, yielding behavior-specific controllers that exhibit limited generalization and brittle performance when deployed on irregular terrains and in diverse situations. To address this challenge, we propose Adaptive Humanoid Control (AHC) that adopts a two-stage framework to learn an adaptive humanoid locomotion controller across different skills and terrains. Specifically, we first train several primary locomotion policies and perform a multi-behavior distillation process to obtain a basic multi-behavior controller, facilitating adaptive behavior switching based on the environment. Then, we perform reinforced fine-tuning by collecting online feedback in performing adaptive behaviors on more diverse terrains, enhancing terrain adaptability for the controller. We conduct experiments in both simulation and real-world experiments in Unitree G1 robots. The results show that our method exhibits strong adaptability across various situations and terrains. Project website: https://ahc-humanoid.github.io.

仿人机器人自适应控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。