arXiv:2604.22226cs.CV2026-04被引 1

让AI看长视频体育比赛时能像人一样找时间证据并推理。

Towards Temporal Compositional Reasoning in Long-Form Sports Videos

论文配图:Towards Temporal Compositional Reasoning in Long-Form Sports Videos
图 1 · 摘自论文原文
  • 用时间奖励机制训练模型,学会定位分散的时间片段
  • 在14000+问题上显著提升时间证据组合与定位准确率
  • 适合研究长视频理解、多模态推理的学者和工程师

体育视频是多模态理解的挑战性领域,因其涉及复杂动态的人类行为。尽管多模态大语言模型(MLLMs)进展迅速,长时程推理仍困难,因回答问题需定位稀疏的时间证据并整合推理。我们归因于两个紧密耦合因素:对分散时间证据的监督不足,以及缺乏要求模型识别、定位并证明时间证据的方法。为此,我们提出SportsTime,一个大规模长时体育视频理解基准,包含14,000+开放式问答对和50,000+逐步时间证据标注。基于SportsTime,我们提出链式时间推理(CoTR),将推理视为时间锚定证据的组合过程。训练中引入时间奖励的GRPO策略以鼓励时间锚定推理;推理时采用锚定-观察-推断的证据搜索循环,迭代定位、验证并组合时间证据后生成答案。实验表明,SportsTime作为基准具有价值,CoTR在多项指标上优于强基线模型,持续提升时间组合推理与步骤级定位质量。

原文摘要 · Abstract (English)

Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in Multimodal Large Language Models (MLLMs), long-horizon reasoning in sports videos remains difficult, as answering questions requires both locating temporally sparse evidence and integrating it into reasoning. We attribute this limitation to two closely coupled factors: insufficient supervision over temporally dispersed evidence, and the lack of methods that require models to identify, localize, and justify temporal evidence. To address these gaps, we introduce SportsTime, a large-scale benchmark for long-form sports video understanding, comprising 14K+ open-ended QA pairs and 50K+ step-wise temporal evidence annotations. Building on SportsTime, we propose Chain-of-Time Reasoning (CoTR), which treats reasoning as a process of temporally grounded evidence composition. Specifically, during training, CoTR introduces a temporal-reward GRPO to encourage temporally grounded reasoning. During inference, it employs an anchor-observe-infer evidence-seeking loop to iteratively localize, verify, and compose temporal evidence before producing the final answer. Experiments demonstrate the usefulness of SportsTime as a benchmark and the effectiveness of CoTR, which consistently improves temporal compositional reasoning and step-wise grounding quality over strong MLLM baselines.

视频理解长视频多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。