用语义引导生成更稳定的视频摘要,提升关键帧选择质量
Semantic-Guided Unsupervised Video Summarization
- 引入帧级语义对齐注意力机制,指导生成器选关键帧
- 在多个数据集上优于现有无监督方法,训练更稳定
- 适合需要高效视频摘要的多模态分析场景
视频摘要对于社交理解至关重要,可帮助高效浏览海量多媒体内容并从社交平台中提取关键信息。现有大多数无监督摘要方法依赖生成对抗网络(GAN)来增强关键帧选择,并通过对抗训练生成连贯的视频摘要。然而,这些方法主要利用单模态特征,忽视了语义信息在关键帧选择中的引导作用,且常面临训练不稳定的问题。为此,我们提出一种新型语义引导的无监督视频摘要方法。具体而言,设计了一种新的帧级语义对齐注意力机制,并将其集成到关键帧选择器中,引导对抗框架内的基于Transformer的生成器更好地重建视频。此外,采用增量训练策略逐步更新模型组件,有效缓解了GAN训练的不稳定性。实验结果表明,该方法在多个基准数据集上均取得更优性能。
原文摘要 · Abstract (English)
Video summarization is a crucial technique for social understanding, enabling efficient browsing of massive multimedia content and extraction of key information from social platforms. Most existing unsupervised summarization methods rely on Generative Adversarial Networks (GANs) to enhance keyframe selection and generate coherent, video summaries through adversarial training. However, such approaches primarily exploit unimodal features, overlooking the guiding role of semantic information in keyframe selection, and often suffer from unstable training. To address these limitations, we propose a novel Semantic-Guided Unsupervised Video Summarization method. Specifically, we design a novel frame-level semantic alignment attention mechanism and integrate it into a keyframe selector, which guides the Transformer-based generator within the adversarial framework to better reconstruct videos. In addition, we adopt an incremental training strategy to progressively update the model components, effectively mitigating the instability of GAN training. Experimental results demonstrate that our approach achieves superior performance on multiple benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。