用真人示范+模拟训练,让机器人几分钟学会打网球反手
TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes

- 将教练-学习者分工显式化,人类提供动作示范与关键窗口
- 单次示范即可零样本泛化到未见目标位置,训练仅需一小时
- 适合快速部署动态技能,尤其适合硬件成本敏感场景
如何学会打网球反手?不是看千场赛事,而是跟教练练习。我们主张这对人形机器人同样有效:动态技能的结果由轨迹中约20cm的关键交互段决定。掌握这一窗口需协调控制、物理和形态。学习因此简化为掌握少数动作并反复练习直至精准。为此提出TaskNPoint训练协议,明确教练-学习者分工:人类提供四类输入——离散技能(如不同击球)、每类一个示范、交互窗口定位、目标任务。在物理真实仿真环境中训练,补全动作轨迹并增强对未建模事件的鲁棒性。训练中随机采样目标位置,使单次示范实现零样本泛化至未见目标。在Unitree G1人形机器人上验证:仅需短时视频示范与单块GPU训练一小时内完成,无需逐任务调奖赏函数,即可完成反手击球、正手击球、踢足球、从新位置抓取放置箱子等任务。
原文摘要 · Abstract (English)
How do we learn to hit a tennis backhand? Not from a thousand hours of tennis tournaments on TV - we work with a coach and practice. We argue this is also the right recipe for teaching dynamic skills to humanoid robots. This follows from a structural property of dynamic skills: the outcome is decided by a short, crucial portion of the trajectory - for a backhand, the ~20cm of racket travel around ball contact. Getting this interaction window right requires coordinating the whole motion, so that control, physics, and morphology act in concert. Learning thus reduces to mastering a handful of distinct actions and, for each, practicing until the window comes out right. To this end, we introduce TaskNPoint, a training protocol which makes the coach-learner division of labor explicit. The human coach contributes four inputs: a discrete set of skills (e.g. different shots), one demonstration per skill, identification of the interaction window, and the goal. Learning in a physically realistic simulation environment fills in each action trajectory and provides robustness to unmodeled events. Crucially, randomized target sampling during training lets a single demonstration generalize zero-shot to unseen goal locations. We test this approach on a Unitree G1 humanoid that hits forehands and backhands against balls thrown by a human, kicks incoming soccer balls, and picks and places boxes from novel locations. We find that learning is successful from short human video demonstrations and under an hour of training on a single GPU, with no per-task reward tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。