arXiv:2504.17771cs.ROcs.AI2025-04ICRA被引 8

混合物理模型与学习策略,让羽毛球机器人更稳更快

Integrating Learning-Based Manipulation and Physics-Based Locomotion for Whole-Body Badminton Robot Control

  • 用物理模型控制底盘,为机械臂提供稳定基础
  • 结合模仿与强化学习,实现94.5%击球成功率
  • 适合研究敏捷移动操作的机器人学者参考

基于学习的方法(如模仿学习和强化学习)能在复杂敏捷任务中生成优秀控制策略,但现有工作尚未将学习策略与基于模型的方法融合以降低训练复杂度并保障安全与稳定性。本文提出哈姆雷特(Hamlet),一种新型混合控制系统,用于敏捷羽毛球机器人。我们设计了一种基于模型的底盘运动策略,为机械臂策略提供基础支持;同时提出一种融入物理信息的“模仿+强化”训练框架,利用带有特权信息的模型化策略,在模仿学习与强化学习阶段引导机械臂策略训练。此外,在模仿学习阶段即训练价值函数,缓解从模仿到强化阶段的性能下降问题。实验在自研羽毛球机器人上进行,对发球机达成94.5%成功率,对人类玩家达90.7%。系统可轻松推广至其他敏捷抓取、乒乓球等任务。项目主页:https://dreamstarring.github.io/HAMLET/

原文摘要 · Abstract (English)

Learning-based methods, such as imitation learning (IL) and reinforcement learning (RL), can produce excel control policies over challenging agile robot tasks, such as sports robot. However, no existing work has harmonized learning-based policy with model-based methods to reduce training complexity and ensure the safety and stability for agile badminton robot control. In this paper, we introduce Hamlet, a novel hybrid control system for agile badminton robots. Specifically, we propose a model-based strategy for chassis locomotion which provides a base for arm policy. We introduce a physics-informed "IL+RL" training framework for learning-based arm policy. In this train framework, a model-based strategy with privileged information is used to guide arm policy training during both IL and RL phases. In addition, we train the critic model during IL phase to alleviate the performance drop issue when transitioning from IL to RL. We present results on our self-engineered badminton robot, achieving 94.5% success rate against the serving machine and 90.7% success rate against human players. Our system can be easily generalized to other agile mobile manipulation tasks such as agile catching and table tennis. Our project website: https://dreamstarring.github.io/HAMLET/.

机器人控制羽毛球机器人混合策略强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。