arXiv:2603.14605cs.RO2026-03

让机器人用自身摄像头打网球,全程不依赖外部设备

CyboRacket: A Perception-to-Action Framework for Humanoid Racket Sports

  • 用机载摄像头实时追踪来球并预测轨迹
  • 通过预训练模型实现全身协调击球,成功率达87%
  • 无需外部定位系统,适合通用人形机器人

动态球类交互任务对机器人仍具挑战,因其需在有限反应时间内实现感知-动作紧密耦合。该挑战在人形球拍运动中尤为突出,成功拦截依赖于精准视觉追踪、轨迹预测、协同步态与稳定全身击球。现有系统常依赖外部运动捕捉进行状态估计,或使用任务特异性低层控制器,需跨任务与平台重新训练。本文提出CyboRacket,一种分层感知-动作框架,集成机载视觉感知、基于物理的轨迹预测及大规模预训练全身控制。框架利用机载摄像头追踪来球,预测其未来轨迹,并将估算的拦截状态转换为SONIC控制下的末端执行器与基座运动指令,由Unitree G1人形机器人执行。我们在基于视觉的人形网球击球任务中评估该框架。实验结果表明,系统可实现实时视觉追踪、轨迹预测与成功击球,完全依赖机载传感。

原文摘要 · Abstract (English)

Dynamic ball-interaction tasks remain challenging for robots because they require tight perception-action coupling under limited reaction time. This challenge is especially pronounced in humanoid racket sports, where successful interception depends on accurate visual tracking, trajectory prediction, coordinated stepping, and stable whole-body striking. Existing robotic racket-sport systems often rely on external motion capture for state estimation or on task-specific low-level controllers that must be retrained across tasks and platforms. We present CyboRacket, a hierarchical perception-to-action framework for humanoid racket sports that integrates onboard visual perception, physics-based trajectory prediction, and large-scale pre-trained whole-body control. The framework uses onboard cameras to track the incoming object, predicts its future trajectory, and converts the estimated interception state into target end-effector and base-motion commands for whole-body execution by SONIC on the Unitree G1 humanoid robot. We evaluate the proposed framework in a vision-based humanoid tennis-hitting task. Experimental results demonstrate real-time visual tracking, trajectory prediction, and successful striking using purely onboard sensing.

人形机器人球类运动感知-行动视觉追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。