arXiv:2504.18689cs.CVcs.AI2025-04被引 2

用分层注意力机制提升教学视频摘要质量,关键步骤更准确。

HierSum: A Global and Local Attention Mechanism for Video Summarization

  • 结合字幕细粒度信息与视频整体指令,分层建模视频结构
  • 以'最常重播'为监督信号,精准识别关键片段,F1得分领先
  • 新构建多模态数据集,提升模型在真实场景的概括能力

视频摘要旨在生成保留关键信息的精简版本。本文聚焦教学类视频,提出一种分层方法HierSum,将视频分解为对应重要步骤的有意义片段。该方法融合字幕提供的细粒度局部线索与视频级指令提供的全局上下文信息,并利用‘最常重播’统计量作为监督信号,定位关键段落,从而提升摘要效果。在TVSum、BLiSS、Mr.HiSum及WikiHow测试集上,HierSum在F1分数和排序相关性等关键指标上持续优于现有方法。此外,我们基于WikiHow和EHow视频及其步骤说明文章构建了一个新的多模态数据集,大量消融实验表明,在该数据集上训练显著提升了模型在目标数据集上的表现。

原文摘要 · Abstract (English)

Video summarization creates an abridged version (i.e., a summary) that provides a quick overview of the video while retaining pertinent information. In this work, we focus on summarizing instructional videos and propose a method for breaking down a video into meaningful segments, each corresponding to essential steps in the video. We propose \textbf{HierSum}, a hierarchical approach that integrates fine-grained local cues from subtitles with global contextual information provided by video-level instructions. Our approach utilizes the ``most replayed" statistic as a supervisory signal to identify critical segments, thereby improving the effectiveness of the summary. We evaluate on benchmark datasets such as TVSum, BLiSS, Mr.HiSum, and the WikiHow test set, and show that HierSum consistently outperforms existing methods in key metrics such as F1-score and rank correlation. We also curate a new multi-modal dataset using WikiHow and EHow videos and associated articles containing step-by-step instructions. Through extensive ablation studies, we demonstrate that training on this dataset significantly enhances summarization on the target datasets.

视频摘要分层注意力教学视频多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。