自动剪辑古典音乐会视频,精准判断切镜时机与画面选择
When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
- 融合音频频谱与视觉特征,用轻量级模型判断何时切镜
- 在多机位视频中准确识别剪辑点,提升画面选择质量
- 适合音乐视频自动化制作,也适用于其他演出场景
自动化视频编辑在计算机视觉与多媒体领域仍属研究空白,尤其相较视频生成与场景理解的热度。本文针对多摄像机录制的古典音乐会视频,将剪辑问题分解为两个子任务:何时切镜与如何切镜。针对时间分割任务,提出一种新型多模态架构,融合音频信号的对数梅尔频谱、可选图像嵌入及标量时间特征,通过轻量级卷积-变压器流水线实现。针对空间选择任务,采用基于CLIP的编码器替代旧有骨干网络(如ResNet),并限制干扰片段仅来自同一场音乐会。数据集通过伪标签方法构建,原始视频被自动聚类为连贯镜头段。实验表明,所提模型在剪辑点检测上优于以往基线,在视觉镜头选择上表现良好,推动了多模态自动化视频编辑的前沿进展。
原文摘要 · Abstract (English)
Automated video editing remains an underexplored task in the computer vision and multimedia domains, especially when contrasted with the growing interest in video generation and scene understanding. In this work, we address the specific challenge of editing multicamera recordings of classical music concerts by decomposing the problem into two key sub-tasks: when to cut and how to cut. Building on recent literature, we propose a novel multimodal architecture for the temporal segmentation task (when to cut), which integrates log-mel spectrograms from the audio signals, plus an optional image embedding, and scalar temporal features through a lightweight convolutional-transformer pipeline. For the spatial selection task (how to cut), we improve the literature by updating from old backbones, e.g. ResNet, with a CLIP-based encoder and constraining distractor selection to segments from the same concert. Our dataset was constructed following a pseudo-labeling approach, in which raw video data was automatically clustered into coherent shot segments. We show that our models outperformed previous baselines in detecting cut points and provide competitive visual shot selection, advancing the state of the art in multimodal automated video editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。