DeepSport用主动推理看懂多运动视频,更准更省帧。
DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
- 通过动态选帧实现边看边思考的主动推理
- 在6.7k数据集上超越顶尖模型,用帧数减少30%以上
- 可零样本迁移至新运动和动作识别任务
体育视频理解需捕捉高速动态、复杂规则与长时序上下文。现有多模态大模型仍局限于单一运动、特定任务或无需训练的范式。我们提出DeepSport,首个端到端训练的多任务、多运动视频理解多模态大模型。它从被动帧处理转向主动迭代推理,动态提取关键帧实现‘与视频共思’。为训练模型,我们通过三步文本-视觉蒸馏流程构建了统一的78,000样本数据集,并采用渐进式两阶段训练策略:先通过体育课程监督微调建立基础感知能力,再结合新型工具使用奖励进行代理强化学习。在涵盖6.7k样本的综合性基准测试中,DeepSport达到领先性能,超越多个商用与开源模型,且显著减少帧使用量。此外,其在未见运动和广泛运动识别任务中表现出强零样本迁移能力,为复杂视频推理建立了高效通用的基础。
原文摘要 · Abstract (English)
Sports video understanding requires perceiving high-speed dynamics, complex rules, and long temporal contexts. Yet, current Multimodal Large Language Models (MLLMs) remain narrowly focused on single sports, specific tasks, or training-free paradigms. We introduce DeepSport, the first end-to-end trained MLLM for multi-task, multi-sport video understanding. DeepSport shifts from passive frame processing to active, iterative reasoning, dynamically extracting frames to "think with videos." To train our model, we curate a unified 78k-sample dataset via a rigorous three-step text-and-vision distillation pipeline. We then employ a progressive two-stage training strategy: a Sports Curriculum Supervised Fine-Tuning phase to build foundational perception, followed by Agentic Reinforcement Learning with a novel tool-use reward. Extensive experiments on a comprehensive 6.7k benchmark demonstrate that DeepSport achieves state-of-the-art performance, outperforming powerful proprietary and open-source models, while utilizing significantly fewer frames. Furthermore, it exhibits strong zero-shot transferability to unseen sports and broad motion recognition tasks, establishing a highly efficient and generalized foundation for complex video reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。