让AI像裁判一样多视角看比赛,提升体育视频理解能力
Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

- 设计智能体框架,主动选择最佳视角并整合多视图证据
- 在1022个比赛片段上测试,性能比最强基线提升15.61%
- 适合研究多视角视觉推理与体育视频分析的学者
当前多模态大模型在单视角视频理解上表现优异,但体育视频存在密集遮挡、快速运动和复杂互动,仅靠单一视角难以解析。现实中赛事由多角度摄像机记录,裁判也依赖多视角互补信息判罚。然而现有基准未评估模型在多视角体育视频上的表现。为此,我们构建了SportMV-Bench,基于官方比赛录像,通过基于LLM生成、MLLM验证与人工筛选的流水线,确保数据质量。该基准包含1022个多视角视频包与3015个问答对,覆盖10项运动,分三类:感知导向识别(PAR)、规则导向事件解析(REI)与判罚决策推理(ADR)。分析表明,当前MLLM无法有效利用多视角信息,瓶颈在于细粒度视觉感知与视角选择,而非逻辑推理或领域知识。为此我们提出SportMV-Agent,一种智能体框架,通过主动视角选择、感知工具执行与证据驱动推理的迭代循环,相较最强基线实现15.61%的相对提升。
原文摘要 · Abstract (English)
Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing benchmark evaluates MLLMs on multi-view sports video understanding. To address this gap, we introduce SportMV-Bench, a comprehensive benchmark built from official match recordings, through a dedicated pipeline combining LLM-based generation, MLLM-based verification, and human filtering to ensure quality and consistency. SportMV-Bench containing 1022 multi-view video bundles and 3015 question-answer pairs spanning 10 sports across three categories: Perception-Aware Recognition (PAR), Rule-aware Event Interpretation (REI), and Adjudicative Decision Reasoning (ADR). Our analysis shows that current MLLMs fail to effectively exploit multi-view information, with the bottlenecks lying in fine-grained visual perception and view selection rather than logical reasoning or domain knowledge. We propose SportMV-Agent, an agentic framework that orchestrates an iterative loop of active view selection, perception tool execution, and evidence-grounded reasoning, achieving a significant 15.61% relative improvement over the strongest MLLM baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。