提出足球视频理解新基准,检验模型是否真看懂画面。
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

- 构建包含13类赛事事件的标注数据集,分三层语义线索评估视觉定位能力。
- 现有顶尖模型在宽松标准下接地率不足50%,且严重忽略时间信息。
- 适合关注视觉推理真实性、模型可解释性的研究者和开发者。
视觉语言模型(VLMs)在足球视频理解中展现出强大潜力。然而,由于足球视频存在视角变化大、镜头切换快、场景杂乱等复杂性,目前尚不清楚这些模型是基于有意义的视觉证据,还是依赖虚假相关性和捷径学习。现有评估主要关注分类准确率,缺乏对视觉定位的衡量。为此,我们提出SoccerLens基准,包含13类常见足球事件的标注视频片段,并将视觉线索按语义相关性分为三个层次。我们扩展Chefer [arXiv:2103.15679]的归因方法,联合建模时空注意力,引入新评估指标以判断模型注意力是否与标注线索对齐,或偏离至虚假区域。对当前领先足球VLMs的评估显示,尽管分类准确率高,但其接地性能在最宽松定义下仍低于50%,且持续低估时间信息。这一结果揭示了预测性能与真实视觉定位之间的巨大差距,凸显在复杂时空领域进行基于定位的评估的必要性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variations, rapid shot transitions, and cluttered scenes, it remains unclear on whether VLMs rely on meaningful visual evidence or exploit spurious correlations and shortcut learning. Existing evaluation protocols focus primarily on classification accuracy and do not assess visual grounding. To address this limitation, we introduce SoccerLens, a benchmark for grounded soccer video understanding. The benchmark contains annotated video segments spanning $13$ common soccer events, with structured visual cues organized into three levels of semantic relevance. We further extend the attribution method of Chefer [arXiv:2103.15679] to jointly model spatial and temporal attention, and introduce evaluation metrics that measure whether model attention aligns with annotated cues or drifts toward spurious regions. Our evaluation of state-of-the-art soccer VLMs shows that, despite strong classification accuracy, current models fail to exceed $50\%$ grounding performance even under the loosest cue definitions and consistently underutilize temporal information. These results reveal a substantial gap between predictive performance and true visual grounding, highlighting the need for grounded evaluation in complex spatio-temporal domains such as soccer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。