用短片段数据提升足球精彩集锦的实时解说生成效果
Commentary Generation for Soccer Highlights
- 基于MatchVoice模型,针对精彩集锦设计细粒度视频-解说对齐方法
- 在GOAL数据集上验证不同窗口大小下零样本性能,发现小窗口更优
- 适合关注体育视频生成与多模态对齐的研究者和开发者
自动足球解说生成已从模板系统发展到先进神经架构,旨在实时描述体育赛事。尽管SoccerNet-Caption奠定了基础,但视频内容与解说之间难以实现细粒度对齐仍是关键挑战。近期工作如MatchTime通过粗粒度与细粒度对齐技术改善了时间同步性。本文将MatchVoice扩展至足球精彩集锦的解说生成,使用强调短片段的GOAL数据集。我们复现了原始MatchTime结果,并评估不同训练配置与硬件限制的影响。此外,研究了不同窗口大小对零样本性能的影响。结果显示,MatchVoice具备良好泛化能力,但需融合更广泛视频-语言领域技术以进一步提升表现。代码开源:https://github.com/chidaksh/SoccerCommentary。
原文摘要 · Abstract (English)
Automated soccer commentary generation has evolved from template-based systems to advanced neural architectures, aiming to produce real-time descriptions of sports events. While frameworks like SoccerNet-Caption laid foundational work, their inability to achieve fine-grained alignment between video content and commentary remains a significant challenge. Recent efforts such as MatchTime, with its MatchVoice model, address this issue through coarse and fine-grained alignment techniques, achieving improved temporal synchronization. In this paper, we extend MatchVoice to commentary generation for soccer highlights using the GOAL dataset, which emphasizes short clips over entire games. We conduct extensive experiments to reproduce the original MatchTime results and evaluate our setup, highlighting the impact of different training configurations and hardware limitations. Furthermore, we explore the effect of varying window sizes on zero-shot performance. While MatchVoice exhibits promising generalization capabilities, our findings suggest the need for integrating techniques from broader video-language domains to further enhance performance. Our code is available at https://github.com/chidaksh/SoccerCommentary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。