arXiv:2409.14543cs.CVcs.AI2024-09被引 16

通过运动注意力图提升网球与羽毛球追踪精度。

TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps

  • 融合视觉特征与可学习的运动注意力图,增强运动目标定位。
  • 在网球和羽毛球数据集上显著提升TrackNetV2/V3性能。
  • 轻量级插件式设计,适合实时体育视频追踪场景。

高动态、小尺寸物体(如体育视频中的球)的精准检测与跟踪面临运动模糊和遮挡等挑战。尽管近期深度学习框架如TrackNetV1、V2和V3已推进网球与羽毛球追踪,但在部分遮挡或低可见性场景下仍表现不佳,主要因模型过度依赖视觉特征而未显式融入运动信息,影响轨迹预测精度。本文提出对TrackNet家族的改进:通过运动感知融合机制,将高层视觉特征与可学习的运动注意力图结合,有效强调运动球体位置,提升追踪性能。该方法利用帧差图,并由运动提示层调制,以突出时间上的关键运动区域。在网球与羽毛球数据集上的实验表明,该方法显著提升了TrackNetV2和V3的追踪效果。我们将其称为基于现有TrackNet架构的轻量级、即插即用解决方案——TrackNetV4。

原文摘要 · Abstract (English)

Accurately detecting and tracking high-speed, small objects, such as balls in sports videos, is challenging due to factors like motion blur and occlusion. Although recent deep learning frameworks like TrackNetV1, V2, and V3 have advanced tennis ball and shuttlecock tracking, they often struggle in scenarios with partial occlusion or low visibility. This is primarily because these models rely heavily on visual features without explicitly incorporating motion information, which is crucial for precise tracking and trajectory prediction. In this paper, we introduce an enhancement to the TrackNet family by fusing high-level visual features with learnable motion attention maps through a motion-aware fusion mechanism, effectively emphasizing the moving ball's location and improving tracking performance. Our approach leverages frame differencing maps, modulated by a motion prompt layer, to highlight key motion regions over time. Experimental results on the tennis ball and shuttlecock datasets show that our method enhances the tracking performance of both TrackNetV2 and V3. We refer to our lightweight, plug-and-play solution, built on top of the existing TrackNet, as TrackNetV4.

目标追踪运动分析注意力机制体育视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。