通过动态切换策略,让机器人在敏捷与稳定间平衡,实现长时序全身控制。
BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control
- 设计在线策略切换框架,融合全局协调与模块化精度两种控制方式。
- 在真实机器人上完成多种长时序运动操作任务,性能超越现有方法。
- 适合需要高鲁棒性与灵活性的复杂机器人控制场景,如人形机器人作业。
尽管控制、强化学习和模仿学习取得了进展,但构建能实现敏捷、精确且鲁棒的全身行为,尤其是在长时序任务中,仍具挑战。现有方法通常采用两种范式:耦合的全身策略用于全局协调,解耦策略用于模块化精度。然而缺乏系统性整合方法,导致敏捷性、鲁棒性与精度之间的权衡未被解决。本文提出BAT框架,一种基于在线策略切换的方法,可动态选择两种互补的全身体强化学习控制器,以在不同运动情境下平衡敏捷性与稳定性。该框架包含两个互补模块:通过分层强化学习并结合滑动窗口策略预评估获得专家引导的切换策略,以及从离散运动标记序列预测选项偏好的选项感知VQ-VAE,以提升泛化能力。最终决策通过两个模块的置信度加权融合得到。大量仿真及在Unitree G1人形机器人上的真实实验表明,BAT能够实现多样化的长时序运动操作任务,并在各类任务中优于先前方法。
原文摘要 · Abstract (English)
Despite recent advances in control, reinforcement learning, and imitation learning, developing a unified framework that can achieve agile, precise, and robust whole-body behaviors, particularly in long-horizon tasks, remains challenging. Existing approaches typically follow two paradigms: coupled whole-body policies for global coordination and decoupled policies for modular precision. However, without a systematic method to integrate both, this trade-off between agility, robustness, and precision remains unresolved. In this work, we propose BAT, an online policy-switching framework that dynamically selects between two complementary whole-body RL controllers to balance agility and stability across different motion contexts. Our framework consists of two complementary modules: a switching policy learned via hierarchical RL with an expert guidance from sliding-horizon policy pre-evaluation, and an option-aware VQ-VAE that predicts option preference from discrete motion token sequences for improved generalization. The final decision is obtained via confidence-weighted fusion of two modules. Extensive simulations and real-world experiments on the Unitree G1 humanoid robot demonstrate that BAT enables versatile long-horizon loco-manipulation and outperforms prior methods across diverse tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。