arXiv:2608.07932cs.CV2026-08

解决体育视频中角色和动作难区分的问题,提升细粒度推理准确率。

SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning

论文配图:SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning
图 1 · 摘自论文原文
  • 用视觉提案+交错定位融合机制,精准定位球员与球的位置。
  • 在两个新数据集上达到顶尖性能,显著优于现有模型。
  • 适合需要高精度动作识别的体育分析与转播系统开发者。

体育视频分析对运动表现评估和转播优化至关重要。密集体育视频推理需在长时序中理解大量小尺度、高度互动且视觉相似的实体(如穿相同队服的球员、球)。当前大模型因缺乏细粒度视觉细节,常依赖文本先验猜测答案,尤其在区分视觉相似的动作和球员时表现不佳。为此,我们提出 SportsGrounder 框架,利用开放词汇视觉专家辅助交错定位,专门应对密集体育视频推理。通过提取领域引导的物体提议,并引入交错定位融合(IGF)机制,逐帧融合显式边界框坐标与隐式视觉语义及全局网格特征,保持严格时间对齐,避免序列长度爆炸。此外,设计动作感知监督(AAS)模块,直接正则化模型隐藏状态,迫使网络学习准确运动表征,而非依赖语言偏见。通过混合偏好优化(MPO)训练,有效区分误导性干扰项。在新构建的密集体育问答数据集(基于 SoccerNet 与 FineSports)上的实验表明,SportsGrounder 显著提升细粒度推理能力,达到当前最优准确率。

原文摘要 · Abstract (English)

Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-scale, highly interactive, and visually homogeneous entities (e.g., players sharing identical uniforms, the ball) across long temporal contexts. Current Large Multimodal Models (LMMs) inherently struggle with such dense visual complexities. Due to the lack of fine-grained visual details, these models often over-rely on textual priors to guess answers, especially when distinguishing visually similar actions and players. To address this, we propose SportsGrounder, a framework that leverages an open-vocabulary visual expert to aid interleaved grounding specifically for dense sports video reasoning. To achieve precise spatial localization, we extract domain-guided object proposals and introduce an Interleaved Grounding Fusion (IGF) mechanism. The IGF frame-by-frame integrates explicit bounding box coordinates and implicit visual semantics with global grid features. This design preserves strict temporal alignment and prevents sequence length explosion. Furthermore, we design an Action-Aware Supervision (AAS) module that directly regularizes the model's hidden states, forcing the network to learn accurate motion representations rather than relying on language bias. Optimized with Mixed Preference Optimization (MPO) to better distinguish deceptive distractors, our extensive experiments on newly curated dense sports VQA datasets (derived from SoccerNet and FineSports) demonstrate that SportsGrounder significantly improves fine-grained reasoning and achieves state-of-the-art accuracy.

体育视频视觉定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。