让音乐随视频节奏与用户指令动态变化,生成更符合预期的配乐。
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
- 分两阶段训练:先对齐视觉与音频时间,再融合多条件控制音乐生成。
- 在主观与客观评测中均优于现有方法,实现更高可控性与同步性。
- 适合需要精准配乐控制的影视创作、短视频制作等场景。
音乐能增强视频叙事与情感表达,推动自动视频到音乐(V2M)生成的需求。然而,现有方法仅依赖视觉特征或附加文本输入,生成过程缺乏透明度,难以满足用户期望。为此,我们提出一种多时变条件引导的V2M生成框架,通过引入多种时间变化的控制条件,提升音乐生成的可控性。方法采用两阶段训练策略:第一阶段设计细粒度特征选择模块与渐进式时间对齐注意力机制,确保视听特征灵活对齐;第二阶段构建动态条件融合模块与控制引导解码器,整合多条件信息并精确指导音乐创作。大量实验表明,该方法在主观和客观评价中均优于现有V2M流水线,显著提升可控性与用户期望的一致性。
原文摘要 · Abstract (English)
Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。