arXiv:2606.06926cs.CVcs.MM2026-06KDD

首个超长体育视频精彩片段检测基准,支持小时级视频分析

SVHighlights: Towards Extremely Long Sport Video Highlight Detection

论文配图:SVHighlights: Towards Extremely Long Sport Video Highlight Detection
图 1 · 摘自论文原文
  • 将长视频分段并融合语义相似镜头,用大模型评估每段重要性
  • 在320段平均2小时的视频上,多项指标领先基线2.5%以上
  • 适合研究长视频理解、体育内容生成与多模态分析的学者

尽管长视频精彩片段检测具有重要应用价值,但现有方法大多局限于短时内容,主要因缺乏合适的基准数据集。为此,我们提出SVHighlights,据我们所知首个面向超长体育视频(单段超1小时)的精彩片段检测基准,覆盖多个体育类别。该数据集通过全片视频与官方精彩剪辑配对构建,采用可扩展的生成管道实现标签标注,无需逐片段打标。数据集包含320个视频,平均时长2.00小时,总计640.18小时,显著超越以往数据集。现有方法在长视频上面临根本挑战:基于短片段训练的模型难以泛化至小时级内容,且片段级评分缺乏全局上下文。为应对这一问题并提供强基线,我们提出TF-SELECTOR——一种无需训练的基于段落的方法,通过合并语义相同的相邻镜头生成上下文感知段落,并利用多模态输入(视觉描述、字幕、音量)的大语言模型预测段落重要性得分。实验表明,相较于经视频时间定位优化的基线,TF-SELECTOR在多数指标上表现更优,提升分别为HIT@1 +2.50,HIT@K +4.04,IoU +2.95。这些结果确立了SVHighlights作为长视频精彩片段检测的挑战性测试平台,并证明简单段落策略可有效扩展至小时级视频。

原文摘要 · Abstract (English)

While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due to the absence of a suitable benchmark. To bridge this gap, we introduce SVHighlights, to the best of our knowledge, the first benchmark for highlight detection in extremely long sports videos, each exceeding one hour in duration, across multiple sports categories. SVHighlights is constructed from pairs of full-length sports videos and their corresponding official highlight videos using a dataset generation pipeline, enabling scalable label generation without conventional per-clip saliency annotation. The benchmark comprises 320 videos with an average duration of 2.00 hours and a total of 640.18 hours, substantially exceeding previous datasets. Existing methods also face fundamental challenges on long videos: models trained on short clips fail to generalize to hour-long content, and their clip-level scoring lacks the broader context needed to identify highlights. To address this and provide a strong baseline, we present TF-SELECTOR, a training-free segment-based approach that divides each video into context-aware segments by merging adjacent shots sharing the same semantic content, and predicts segment-level saliency scores using a large language model with multimodal inputs including visual captions, transcripts, and audio volume. Experiments demonstrate that TF-SELECTOR achieves superior performance across most metrics compared to Video Temporal Grounding (VTG)-tuned baselines, with improvements of +2.50 in HIT@1, +4.04 in HIT@K, and +2.95 in IoU. These results establish SVHighlights as a challenging testbed for long-form highlight detection and demonstrate that a simple segment-based strategy can effectively scale to hour-long videos.

视频生成长视频多模态体育分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。