用扩散模型生成多种符合不同人偏好的视频摘要
SummDiff: Generative Modeling of Video Summarization with Diffusion
- 将视频摘要视为条件生成任务,学习摘要的分布
- 在多个基准上达到顶尖性能,且贴近个体标注者偏好
- 首次引入扩散模型并分析被忽略的背包问题
视频摘要任务旨在通过选择关键帧来缩短视频,同时保留其核心内容。尽管该任务具有固有的主观性,以往方法通常对多名评分者的平均帧得分进行确定性回归,忽略了好摘要的主观差异。本文提出新范式:将视频摘要建模为条件生成任务,使模型能够学习优质摘要的分布,并生成多个符合不同人类视角的合理摘要。首次在视频摘要中采用扩散模型,SummDiff能动态适应视觉上下文,生成多个以输入视频为条件的候选摘要。大量实验表明,SummDiff不仅在多个基准上达到最先进水平,生成的摘要也更贴近个体标注者偏好。此外,我们通过分析重要的背包问题(knapsack),提出了新颖评估指标,揭示了以往评估中被忽视的关键环节。
原文摘要 · Abstract (English)
Video summarization is a task of shortening a video by choosing a subset of frames while preserving its essential moments. Despite the innate subjectivity of the task, previous works have deterministically regressed to an averaged frame score over multiple raters, ignoring the inherent subjectivity of what constitutes a good summary. We propose a novel problem formulation by framing video summarization as a conditional generation task, allowing a model to learn the distribution of good summaries and to generate multiple plausible summaries that better reflect varying human perspectives. Adopting diffusion models for the first time in video summarization, our proposed method, SummDiff, dynamically adapts to visual contexts and generates multiple candidate summaries conditioned on the input video. Extensive experiments demonstrate that SummDiff not only achieves the state-of-the-art performance on various benchmarks but also produces summaries that closely align with individual annotator preferences. Moreover, we provide a deeper insight with novel metrics from an analysis of the knapsack, which is an important last step of generating summaries but has been overlooked in evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。