arXiv:2606.11092cs.ROcs.AI2026-06

用分阶段强化学习让机器人踢球更准更稳,实测接近职业球员水平。

RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning

论文配图:RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning
图 1 · 摘自论文原文
  • 分三阶段渐进训练:先学稳定动作,再练定点射门,最后实现移动中射门。
  • 仿真中射门误差降低48.6%,球速达13.10米/秒,接近职业水平。
  • 适合需要高精度人形机器人运动控制的研究与工程应用。

顶尖人形机器人射门需全身稳定、高冲击协同动作及精准瞄准。基于运动追踪的强化学习可提升全身协调性,但固定参考难适应不同球位和击球时机;而仅靠任务奖励的强化学习则难以从零探索有效踢法。为此,我们提出RoboNaldo,一种三阶段运动引导课程强化学习框架,用于高冲击人形交互。以单次人类踢球动作为基准,逐步优化为射门性能。课程首先学习稳定全身踢球策略,随后在球静止于随机位置的任意球场景中适配踢法,最后通过运动指令与踢球触发接口扩展至移动球射门。训练期间由高层启发式规划器控制该接口,推理时可换用其他高层控制器驱动同一底层策略。仿真结果显示,相比基线方法,自由球射门误差降低48.6%,球速提升2.96倍。真实世界测试中,在Unitree G1机器人上搭载本地感知系统,自由球与移动球射门平均误差分别为0.73米和0.86米(距目标3米),后着球速达13.10米/秒,相当于专业比赛中开球速度的59-71%。项目页面:https://opendrivelab.com/RoboNaldo。

原文摘要 · Abstract (English)

Elite humanoid soccer shooting requires whole-body stability, high-impulse whole-body interactions, and accuracy to targets. Motion tracking-driven reinforcement learning (RL) provides stability in whole-body movement coordination, but a fixed reference makes it hard to adapt to varied ball positions and strike timings; in contrast, task reward-driven RL struggles to explore and discover valid kicks from scratch. We therefore introduce RoboNaldo, a three-stage motion-guided curriculum RL framework for high-impulse humanoid interaction. A single human-kick reference is used as a scaffold and progressively shifts optimization towards shooting performance. The curriculum first learns a stable whole-body kicking prior, then adapts the kick to free-kick settings where the ball is stationary at random positions, and finally extends it to moving-ball shooting through a locomotion-command and kick-trigger interface. A high-level heuristic planner controls this interface during training, while alternative high-level controllers can drive the same low-level policy at inference. In simulation, RoboNaldo demonstrates free-kick shot error 48.6% lower and shoot velocity 2.96x than prior work baselines. In real world on a Unitree G1 with onboard perception, RoboNaldo attains 0.73 m and 0.86 m average target shooting error from 3 m away in free-kick and moving-ball cases, accordingly. And the post-contact ball velocity reaches 13.10 m/s, which is 59-71% of reported professional open-play shot speed. Project page: https://opendrivelab.com/RoboNaldo.

人形机器人强化学习足球机器人运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。