arXiv:2509.21690cs.RO2025-09被引 7

让机器人用物理预测提前规划击球动作,实现高精度人形乒乓球对打。

PACE: Physics Augmentation for Coordinated End-to-end Reinforcement Learning toward Versatile Humanoid Table Tennis

  • 用球的运动轨迹预测增强观察,让机器人提前决策。
  • 仿真中发球范围广,击球成功率超92%,命中率超96%。
  • 零样本部署到真实机器人,能自动协调脚步与挥拍。

人形乒乓球需要快速感知、主动全身运动和敏捷步法,而端到端控制策略仍难实现。本文提出一种强化学习框架,将球的位置直接映射为四肢关节指令,通过预测信号和密集的物理引导奖励进行增强。一个轻量级学习预测器利用近期球位估计未来状态,为策略提供前瞻性观测。训练时,基于物理的预测器生成精确未来状态,构建密集、信息丰富的奖励以促进有效探索。所获策略在仿真中对多种发球范围均表现优异(命中率≥96%,成功率达≥92%)。消融实验表明,学习预测器和预测奖励设计对端到端学习至关重要。该策略零样本部署于具23个旋转关节的真实机器人Booster T1上,实现精准快速回击,并协调横向与前后步法,展示出通往通用竞技级人形乒乓球的实际路径。代码已开源:https://github.com/purdue-tracelab/TTRL-ICRA2026。

原文摘要 · Abstract (English)

Humanoid table tennis (TT) demands rapid perception, proactive whole-body motion, and agile footwork under strict timing--capabilities that remain difficult for end-to-end control policies. We propose a reinforcement learning (RL) framework that maps ball-position observations directly to whole-body joint commands for both arm striking and leg locomotion, strengthened by predictive signals and dense, physics-guided rewards. A lightweight learned predictor, fed with recent ball positions, estimates future ball states and augments the policy's observations for proactive decision-making. During training, a physics-based predictor supplies precise future states to construct dense, informative rewards that lead to effective exploration. The resulting policy attains strong performance across varied serve ranges (hit rate$\geq$96% and success rate$\geq$92%) in simulations. Ablation studies confirm that both the learned predictor and the predictive reward design are critical for end-to-end learning. Deployed zero-shot on a physical Booster T1 humanoid with 23 revolute joints, the policy produces coordinated lateral and forward-backward footwork with accurate, fast returns, suggesting a practical path toward versatile, competitive humanoid TT. We have open-sourced our RL training code at: https://github.com/purdue-tracelab/TTRL-ICRA2026

人形机器人强化学习乒乓球物理预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。