arXiv:2604.01383cs.CVcs.AI2026-04中稿 · CVPR被引 1

无需训练即可精准定位美式橄榄球接触起始帧,解决真实场景下的动作定位难题。

GRAZE: Grounded Refinement and Motion-Aware Zero-Shot Event Localization

  • 用视觉模型发现候选接触点,再通过运动感知时间推理精炼
  • 在738段视频中97.4%成功输出,±10帧内定位率达77.5%
  • 适合无标注数据、复杂场景下的生物力学分析应用

美式橄榄球训练视频规模大,但目标交互仅占片段中极短时间。可靠生物力学分析依赖于对交互实体及接触起始时刻的精确时空定位。本文研究首次接触帧(FPOC),即球员首次触碰挡靶的帧,在存在相机运动、背景杂乱、多名相似运动员和快速姿态变化的真实训练视频中的定位问题。提出GRAZE——一种无需训练的端到端定位流程,不需任何标注接触样本。GRAZE利用Grounding DINO发现候选球员-挡靶交互,通过运动感知的时间推理进行精炼,并采用SAM2作为像素级接触验证器,而非依赖检测置信度。该发现与验证分离的设计提升了对杂乱场景和接触瞬间不稳定定位的鲁棒性。在738段挡靶训练视频上,97.4%的视频产生有效输出,其中77.5%的视频定位在±10帧范围内,82.7%在±20帧范围内。结果表明,无需任务特定训练即可实现真实场景下帧级接触起始定位。

原文摘要 · Abstract (English)

American football practice generates video at scale, yet the interaction of interest occupies only a brief window of each long, untrimmed clip. Reliable biomechanical analysis, therefore, depends on spatiotemporal localization that identifies both the interacting entities and the onset of contact. We study First Point of Contact (FPOC), defined as the first frame in which a player physically touches a tackle dummy, in unconstrained practice footage with camera motion, clutter, multiple similarly equipped athletes, and rapid pose changes around impact. We present GRAZE, a training-free pipeline for FPOC localization that requires no labeled tackle-contact examples. GRAZE uses Grounding DINO to discover candidate player-dummy interactions, refines them with motion-aware temporal reasoning, and uses SAM2 as an explicit pixel-level verifier of contact rather than relying on detection confidence alone. This separation between candidate discovery and contact confirmation makes the approach robust to cluttered scenes and unstable grounding near impact. On 738 tackle-practice videos, GRAZE produces valid outputs for 97.4% of clips and localizes FPOC within $\pm$ 10 frames on 77.5% of all clips and within $\pm$ 20 frames on 82.7% of all clips. These results show that frame-accurate contact onset localization in real-world practice footage is feasible without task-specific training.

动作定位零样本视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。