用1000场乒乓球对打数据,实现毫秒级事件相机球体状态实时估计。
1000 Rallies: An Event-Camera Dataset and Real-Time Learned Ball-State Estimation for Robotic Table Tennis

- 基于事件相机与高速摄像机同步采集,构建高精度球体运动数据集
- 神经网络联合估计球的位置与速度,使击球点预测误差降低36%
- 首次实现实时事件感知驱动机器人与人对打,适合高速动态场景研究
机器人乒乓球已成为实时感知的重要基准任务,因其球速快、时间要求严苛。准确、高频且低延迟的球体状态估计对轨迹预测与及时控制至关重要。传统帧式相机存在固有权衡:低帧率会遗漏快速移动物体,高帧率则增加数据与计算成本。事件相机则具备微秒级时间分辨率,在充足光照下几乎无运动模糊,适用于高速场景。然而,社区缺乏真实体育场景下的大规模事件感知数据集。本文提出首个大型事件相机乒乓球数据集,包含超过1000场来自业余至顶尖选手的对打录像。每段视频同步记录事件流与14路200 FPS高速帧式相机数据,用于生成1 kHz伪真值标签(球位置、速度、旋转)。基于此数据集,训练了一个对背景人体运动鲁棒的卷积神经网络,从事件中联合估计图像平面内的球位置与速度。将预测速度作为卡尔曼滤波的额外测量,使击球点预测误差相比仅使用位置的基线降低36%。最后,将事件感知系统集成至Stäubli机械臂,首次实现由事件感知驱动的真人-机器人实时乒乓球对打。
原文摘要 · Abstract (English)
Robotic table tennis has emerged as a compelling benchmark for real-time robotic perception due to its fast ball dynamics and stringent timing requirements. Accurate, high-frequency, and low-latency ball state estimation is critical for reliable trajectory prediction and timely control. Traditional frame-based cameras face an inherent trade-off: low frame rates leave temporal blind spots that miss fast-moving objects and high frame rates raise data and computational cost. Event cameras instead offer microsecond temporal resolution and, under sufficient illumination, remain largely free of motion blur even at high ball speeds. However, the community lacks large-scale datasets to develop and benchmark event-based perception in realistic sports scenarios. We address this gap by introducing the first large-scale event-camera dataset for table tennis, comprising over 1000 rallies from a diverse group of players ranging from amateurs to elite-level athletes. Each recording captures the event stream alongside 14 synchronized high-speed frame-based cameras at 200 FPS, which we use to produce 1 kHz pseudo ground-truth labels for ball position, velocity, and spin. Building on this dataset, we train a convolutional neural network robust to background player motion that jointly estimates the ball's position and velocity in the image-plane from events. Treating the predicted velocity as an additional measurement in the Kalman filter reduces bounce-point prediction error by 36% relative to a position-only baseline. Finally, we close the perception-action loop by integrating the event-based system with a Stäubli robotic arm, enabling the first real-time human-robot table tennis rallies driven by event-based perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。