用预训练视觉语言模型+强化学习,自动剪辑电影类通用视频。
A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language Model
- 用预训练视觉语言模型提取编辑上下文特征,替代特定场景特征。
- 在真实电影数据集上验证,生成视频更接近专业剪辑效果。
- 适合想做通用视频自动剪辑的研究者和开发者。
在视频时代,自动视频编辑技术因能减轻工作量、降低对人工编辑的要求而受到产业界和学术界的广泛关注。现有自动编辑系统多针对特定场景(如足球赛事转播),而针对涵盖多种场景与事件的通用编辑任务(如电影或vlog剪辑)研究较少,且将事件驱动的编辑方法拓展至通用场景极具挑战。本文提出一种两阶段通用编辑方案:首先,不同于以往提取特定场景特征的方法,采用预训练视觉语言模型(VLM)提取与编辑相关的上下文表示;其次,为缩小专业视频与自动生成视频之间的差距,设计基于强化学习(RL)的编辑框架,将编辑问题建模为序列决策过程,并训练虚拟编辑器做出更优的连续剪辑决策。最后,在一个更具通用性的电影数据集上评估所提方法。实验结果表明,所提出的上下文表示及强化学习框架具有显著有效性与优势。
原文摘要 · Abstract (English)
In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly scene- or event-specific, e.g., soccer game broadcasting, yet the automatic systems for general editing, e.g., movie or vlog editing which covers various scenes and events, were rarely studied before, and converting the event-driven editing method to a general scene is nontrivial. In this paper, we propose a two-stage scheme for general editing. Firstly, unlike previous works that extract scene-specific features, we leverage the pre-trained Vision-Language Model (VLM) to extract the editing-relevant representations as editing context. Moreover, to close the gap between the professional-looking videos and the automatic productions generated with simple guidelines, we propose a Reinforcement Learning (RL)-based editing framework to formulate the editing problem and train the virtual editor to make better sequential editing decisions. Finally, we evaluate the proposed method on a more general editing task with a real movie dataset. Experimental results demonstrate the effectiveness and benefits of the proposed context representation and the learning ability of our RL-based editing framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。