用AI自动剪辑固定摄像头拍摄的全景视频,让静态画面像电影一样生动。
EditIQ: Automated Cinematic Editing of Static Wide-Angle Videos via Dialogue Interpretation and Saliency Cues
- 通过虚拟摄像机生成多视角画面,模拟真人摄制。
- 结合对话理解与视觉显著性,选出最精彩镜头组合。
- 适合需要自动影视化处理的演出、活动等场景。
我们提出 EditIQ,一个完全自动化的框架,用于对静止广角高分辨率摄像头捕捉的场景进行电影级剪辑。从静态视频流中,EditIQ 首先生成多个虚拟视频流,模拟一组摄像师的视角,这些虚拟镜头称为 'rushes'。随后,通过自动化剪辑算法将这些镜头组装成最能呈现画面张力的影片。为理解关键场景元素并指导剪辑过程,采用双路径方法:(1) 基于大语言模型(LLM)的对话理解模块分析对话流程;(2) 视觉显著性预测识别有意义的场景元素及对应镜头。我们将电影剪辑建模为基于镜头选择的能量最小化问题,其中电影约束决定镜头选择、转场和连贯性。EditIQ 在保持电影一致性和流畅观感的同时,合成出原叙事的美学且视觉吸引人的表现形式。在 BBC Old School 数据集和 11 个剧场表演视频上,通过 20 名参与者的心理物理学实验验证了其优于现有基线的效果。视频样例可访问 https://editiq-ave.github.io/。
原文摘要 · Abstract (English)
We present EditIQ, a completely automated framework for cinematically editing scenes captured via a stationary, large field-of-view and high-resolution camera. From the static camera feed, EditIQ initially generates multiple virtual feeds, emulating a team of cameramen. These virtual camera shots termed rushes are subsequently assembled using an automated editing algorithm, whose objective is to present the viewer with the most vivid scene content. To understand key scene elements and guide the editing process, we employ a two-pronged approach: (1) a large language model (LLM)-based dialogue understanding module to analyze conversational flow, coupled with (2) visual saliency prediction to identify meaningful scene elements and camera shots therefrom. We then formulate cinematic video editing as an energy minimization problem over shot selection, where cinematic constraints determine shot choices, transitions, and continuity. EditIQ synthesizes an aesthetically and visually compelling representation of the original narrative while maintaining cinematic coherence and a smooth viewing experience. Efficacy of EditIQ against competing baselines is demonstrated via a psychophysical study involving twenty participants on the BBC Old School dataset plus eleven theatre performance videos. Video samples from EditIQ can be found at https://editiq-ave.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。