首个端到端足球解说生成模型,实现全程视频精准时间定位与语义描述。
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
- 端到端联合预测时间戳与生成解说,支持45分钟比赛全局建模。
- 在全场比赛视频上达成当前最佳性能,时间对齐准确、语义相关性强。
- 创新帧压缩模块提升长视频理解能力,适合体育视频生成研究者使用。
足球是全球广受欢迎的体育赛事,通常具有长时间的比赛和独特的精彩瞬间。近年来,多模态大语言模型(MLLM)在时间定位和视频理解方面展现出强大潜力,但现有足球解说生成方法多依赖预设时间先验,无法实现端到端处理;传统两阶段方法复杂且难以捕捉全局上下文,性能受限。为此,我们提出TimeSoccer,首个面向完整比赛视频的单锚点密集视频字幕生成(SDVC)端到端足球多模态大模型。TimeSoccer在单次推理中联合预测时间戳并生成解说,实现45分钟比赛的全局上下文建模。为增强长视频理解能力,引入无需训练的运动感知帧压缩模块MoFA-Select,通过粗到精策略自适应选择代表性帧,并融合互补训练范式强化模型对长时序序列的处理能力。大量实验表明,TimeSoccer在端到端框架下达到当前最优(SoTA)性能,生成高质量解说,具备精确的时间对齐与强语义相关性。
原文摘要 · Abstract (English)
Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal grounding and video understanding, soccer commentary generation often requires precise temporal localization and semantically rich descriptions over long-form video. However, existing soccer MLLMs often rely on the temporal a priori for caption generation, so they cannot process the soccer video end-to-end. While some traditional approaches follow a two-step paradigm that is complex and fails to capture the global context to achieve suboptimal performance. To solve the above issues, we present TimeSoccer, the first end-to-end soccer MLLM for Single-anchor Dense Video Captioning (SDVC) in full-match soccer videos. TimeSoccer jointly predicts timestamps and generates captions in a single pass, enabling global context modeling across 45-minute matches. To support long video understanding of soccer matches, we introduce MoFA-Select, a training-free, motion-aware frame compression module that adaptively selects representative frames via a coarse-to-fine strategy, and incorporates complementary training paradigms to strengthen the model's ability to handle long temporal sequences. Extensive experiments demonstrate that our TimeSoccer achieves State-of-The-Art (SoTA) performance on the SDVC task in an end-to-end form, generating high-quality commentary with accurate temporal alignment and strong semantic relevance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。