用强化学习让无人机用球拍连续击球,真实世界平均311次,远超传统方法。
JuggleRL: Mastering Ball Juggling with a Quadrotor via Deep Reinforcement Learning
- 通过大规模仿真训练,结合动态校准与奖励函数设计,提升策略泛化能力。
- 真实世界平均击球311次,最高达462次,显著超越基于模型的基线方法。
- 零样本部署到真实硬件,可适应不同重量球体,适合复杂交互任务研究者。
空中机器人与物体交互需在不确定环境下完成精准、高接触频率的操作。本文研究了配备球拍的四旋翼无人机进行空中击球的任务,该任务要求精确的时间控制、稳定控制及持续适应能力。提出JuggleRL,首个基于强化学习的空中击球系统。在大规模仿真中通过系统性校准四旋翼与球体动力学以减小仿真到现实的差距。训练中采用奖励塑形鼓励以球拍为中心的击球和持续击球,并通过球位置与恢复系数的领域随机化增强鲁棒性与可迁移性。所学策略输出中层指令,由底层控制器执行,实现零样本部署至真实硬件;改进的感知模块结合轻量通信协议,降低高频状态估计延迟,确保实时控制。实验表明,JuggleRL在真实世界连续10次试验中平均达311次击球,最高462次,远超基于模型的基线(最多14次,平均3.1次)。此外,策略能泛化至未见条件,成功击打5克轻球,平均达145.9次击球。本工作证明强化学习可赋予空中机器人在动态交互任务中的鲁棒稳定控制能力。
原文摘要 · Abstract (English)
Aerial robots interacting with objects must perform precise, contact-rich maneuvers under uncertainty. In this paper, we study the problem of aerial ball juggling using a quadrotor equipped with a racket, a task that demands accurate timing, stable control, and continuous adaptation. We propose JuggleRL, the first reinforcement learning-based system for aerial juggling. It learns closed-loop policies in large-scale simulation using systematic calibration of quadrotor and ball dynamics to reduce the sim-to-real gap. The training incorporates reward shaping to encourage racket-centered hits and sustained juggling, as well as domain randomization over ball position and coefficient of restitution to enhance robustness and transferability. The learned policy outputs mid-level commands executed by a low-level controller and is deployed zero-shot on real hardware, where an enhanced perception module with a lightweight communication protocol reduces delays in high-frequency state estimation and ensures real-time control. Experiments show that JuggleRL achieves an average of $311$ hits over $10$ consecutive trials in the real world, with a maximum of $462$ hits observed, far exceeding a model-based baseline that reaches at most $14$ hits with an average of $3.1$. Moreover, the policy generalizes to unseen conditions, successfully juggling a lighter $5$ g ball with an average of $145.9$ hits. This work demonstrates that reinforcement learning can empower aerial robots with robust and stable control in dynamic interaction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。