arXiv:2511.02505cs.CVcs.AI2025-11被引 2

用能量模型优化镜头组合,让自动剪辑更贴近导演风格。

ESA: Energy-Based Shot Assembly Optimization for Automatic Video Editing

  • 基于能量模型学习参考视频的镜头编排风格
  • 通过语义匹配筛选候选镜头,再按风格评分排序
  • 新手也能生成符合叙事逻辑的流畅视频

镜头组装是影视制作与视频编辑中的关键步骤,涉及镜头的排序与排列以构建叙事、传递信息或引发情感。传统上由经验丰富的编辑手动完成。尽管当前智能视频编辑技术可处理部分自动化任务,但往往难以捕捉创作者独特的艺术表达。为此,我们提出一种基于能量函数的镜头组装优化方法:首先利用大语言模型生成剧本,并与视频库进行视觉-语义匹配,获取与剧本语义对齐的候选镜头子集;接着对参考视频中的镜头进行分割与标注,提取景别、摄像机运动、语义等属性;然后使用能量模型学习这些属性,根据参考风格对候选镜头序列进行评分;最后结合多种语法规则,实现镜头组装优化,生成与参考视频风格一致的视频。该方法不仅可依据特定逻辑、叙事需求或艺术风格自动排列组合独立镜头,还能学习参考视频的编排风格,生成连贯的视觉序列或整体视觉表达。即使无视频编辑经验的用户,也能创作出视觉吸引人的视频。

原文摘要 · Abstract (English)

Shot assembly is a crucial step in film production and video editing, involving the sequencing and arrangement of shots to construct a narrative, convey information, or evoke emotions. Traditionally, this process has been manually executed by experienced editors. While current intelligent video editing technologies can handle some automated video editing tasks, they often fail to capture the creator's unique artistic expression in shot assembly. To address this challenge, we propose an energy-based optimization method for video shot assembly. Specifically, we first perform visual-semantic matching between the script generated by a large language model and a video library to obtain subsets of candidate shots aligned with the script semantics. Next, we segment and label the shots from reference videos, extracting attributes such as shot size, camera motion, and semantics. We then employ energy-based models to learn from these attributes, scoring candidate shot sequences based on their alignment with reference styles. Finally, we achieve shot assembly optimization by combining multiple syntax rules, producing videos that align with the assembly style of the reference videos. Our method not only automates the arrangement and combination of independent shots according to specific logic, narrative requirements, or artistic styles but also learns the assembly style of reference videos, creating a coherent visual sequence or holistic visual expression. With our system, even users with no prior video editing experience can create visually compelling videos. Project page: https://sobeymil.github.io/esa.com

自动剪辑能量模型镜头排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。