arXiv:2512.06571cs.RO2025-12被引 7

让机器人在感官混乱下稳定踢球,实现真实场景下的持续精准射门。

Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input

  • 分四阶段训练:先学追球、再学踢球,最后学生模仿教师并适应噪声
  • 在仿真与真实机器人上均实现高精度踢球,多种球门配置下成功率超90%
  • 关键创新是加入噪声建模和在线约束强化学习,提升感知不确定性下的鲁棒性

学习快速且稳定的踢球技能是类人足球机器人的核心能力,但因需快速摆腿、单脚支撑平衡,以及在传感器噪声和外部干扰(如对手)下的鲁棒性挑战而难以实现。本文提出一种基于强化学习的系统,使类人机器人能在不同球-门配置下持续稳健地踢球。系统扩展了典型的教师-学生训练框架:教师使用真实状态信息训练,学生则在含噪声的不完整感知下模仿学习。训练分为四个阶段:(1) 远距离追球(教师);(2) 方向性踢球(教师);(3) 教师策略蒸馏(学生);(4) 学生自适应与优化(学生)。关键设计包括定制奖励函数、真实噪声建模及在线约束强化学习,有效缩小仿真到现实的差距,并在感知不确定下维持性能。仿真与真实机器人上的大量评估显示,在多种球门配置下踢球准确率高,进球成功率超过90%。消融实验进一步证明约束强化学习、噪声建模和自适应阶段的必要性。该工作建立了一个在感知不完善条件下学习类人全身控制视觉运动技能的基准任务。

原文摘要 · Abstract (English)

Learning fast and robust ball-kicking skills is a critical capability for humanoid soccer robots, yet it remains a challenging problem due to the need for rapid leg swings, postural stability on a single support foot, and robustness under noisy sensory input and external perturbations (e.g., opponents). This paper presents a reinforcement learning (RL)-based system that enables humanoid robots to execute robust continual ball-kicking with adaptability to different ball-goal configurations. The system extends a typical teacher-student training framework -- in which a "teacher" policy is trained with ground truth state information and the "student" learns to mimic it with noisy, imperfect sensing -- by including four training stages: (1) long-distance ball chasing (teacher); (2) directional kicking (teacher); (3) teacher policy distillation (student); and (4) student adaptation and refinement (student). Key design elements -- including tailored reward functions, realistic noise modeling, and online constrained RL for adaptation and refinement -- are critical for closing the sim-to-real gap and sustaining performance under perceptual uncertainty. Extensive evaluations in both simulation and on a real robot demonstrate strong kicking accuracy and goal-scoring success across diverse ball-goal configurations. Ablation studies further highlight the necessity of the constrained RL, noise modeling, and the adaptation stage. This work presents a system for learning robust continual humanoid ball-kicking under imperfect perception, establishing a benchmark task for visuomotor skill learning in humanoid whole-body control.

类人机器人强化学习踢球技能感知鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。