arXiv:2509.01439cs.CVcs.AI2025-09中稿 · MMSports 2025被引 2

构建首个足球视频摘要基准数据集,助力自动精彩集锦生成

SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization

  • 基于3国联赛237场广播视频,标注镜头边界构建数据集
  • 提出专用模型达0.3956的测试集F1分数,提升摘要质量
  • 新增长度约束评估指标,更客观衡量生成内容

视频摘要旨在从长视频中提取关键片段,生成简洁且信息丰富的摘要。在体育领域,精彩集锦能捕捉比赛的重要时刻、显著反应及特定情境事件。自动摘要生成可帮助体育媒体从业者减少识别关键片段的时间与精力。然而,缺乏公开可用的数据集制约了运动精彩集锦生成模型的发展。本文通过引入一个专为足球视频摘要设计的精选数据集,填补这一空白。该数据集包含来自西班牙、法国和意大利联赛的237场比赛的镜头边界信息,数据源自SoccerNet数据集。同时,我们提出了一个针对该任务的基线模型,在测试集上取得0.3956的F1分数。此外,我们还设计了一种受目标摘要长度约束的新评估指标,实现更客观的内容评估。数据集与代码已开源:https://ipcv.github.io/SoccerHigh/

原文摘要 · Abstract (English)

Video summarization aims to extract key shots from longer videos to produce concise and informative summaries. One of its most common applications is in sports, where highlight reels capture the most important moments of a game, along with notable reactions and specific contextual events. Automatic summary generation can support video editors in the sports media industry by reducing the time and effort required to identify key segments. However, the lack of publicly available datasets poses a challenge in developing robust models for sports highlight generation. In this paper, we address this gap by introducing a curated dataset for soccer video summarization, designed to serve as a benchmark for the task. The dataset includes shot boundaries for 237 matches from the Spanish, French, and Italian leagues, using broadcast footage sourced from the SoccerNet dataset. Alongside the dataset, we propose a baseline model specifically designed for this task, which achieves an F1 score of 0.3956 in the test set. Furthermore, we propose a new metric constrained by the length of each target summary, enabling a more objective evaluation of the generated content. The dataset and code are available at https://ipcv.github.io/SoccerHigh/.

视频摘要足球分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。