让电影预告片剪辑随音乐节奏弹性变化,更符合专业剪辑规律。
BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation

- 用音乐-视觉对齐编码器和动态规划算法实现弹性剪辑
- 在多个评估维度上超越现有方法,端到端生成完整预告片
- 适合影视自动化生成、创意智能等场景
自动电影预告片生成需从长片中选取镜头并同步背景音乐。现有方法或把音乐对齐放在后期处理,或强制执行一对一的镜头-音乐映射,忽视了专业剪辑节奏具有弹性:快速剪辑对应高能量段落,而平稳镜头则覆盖较安静的乐句。本文提出BEAT框架,包含两个核心组件:MuVA——一种基于Sinkhorn正则化两阶段训练的紧凑音乐-视觉对齐编码器;以及Bar-DP——一种能量自适应的动态规划算法,能根据音乐动态生成弹性多对一的对齐关系。二者集成于五阶段智能体流水线中,以学习到的跨模态特征为基础进行对齐,并通过结构化文本信号协调高层创作决策。为支持全面评估,我们还引入TrailerArena基准,涵盖20+指标,覆盖四个互补维度。在TrailerArena上,BEAT在镜头选择、排序及感知质量方面均达到当前最优表现,且可端到端生成完整预告片。
原文摘要 · Abstract (English)
Automatic movie trailer generation must select shots from a full-length film and synchronize them with background music. Existing methods either relegate music alignment to post-processing or enforce rigid one-to-one shot-music mappings, overlooking that professional editing rhythm is elastic: rapid cuts accompany high-energy passages while sustained shots span quieter bars. We introduce BEAT, a framework that addresses this gap with two core components: MuVA, a compact music-visual alignment encoder trained with Sinkhorn-regularized two-stage learning, and Bar-DP, an energy-adaptive dynamic programming algorithm that produces elastic many-to-one alignments following musical dynamics. These components are integrated into a five-phase agentic pipeline that grounds the core alignment in learned cross-modal features while coordinating higher-level creative decisions through structured text signals. To support comprehensive evaluation, we also introduce TrailerArena, a benchmark with 20+ metrics across four complementary dimensions. On TrailerArena, BEAT achieves state-of-the-art performance across shot selection, ordering, and perceptual quality, while producing fully composed trailers end-to-end.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。