arXiv:2603.11042cs.CVcs.AI2026-03被引 2

无需配对数据,让音乐精准匹配视频节奏和情绪变化。

V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

  • 用视频与音乐各自内部变化规律生成时间曲线,实现跨模态对齐。
  • 在多个数据集上音质提升5-9%,时间对齐度提高21-52%。
  • 支持独立控制音乐风格与时间节奏,适合影视配乐等场景。

现有文本到音乐模型难以实现视频事件与音乐的精细时间对齐。本文提出V2M-ZERO,一种无需视频-音乐配对数据的视频到音乐生成方法,能分离控制时间同步与语义特征(如流派、情绪)。核心思路是:时间对齐只需匹配变化的时机与程度,而非内容本身。尽管音乐与视觉事件语义不同,但具有共享的时间结构,可通过预训练编码器计算模态内相似性得到事件曲线,独立捕捉各模态的时间变化特征,实现跨模态可比表示。训练时仅需微调文本到音乐模型以适应音乐事件曲线,推理时替换为视频事件曲线,无需跨模态训练或成对数据。在OES-Pub、MovieGenBench-Music和AIST++上均达到当前最佳性能,音频质量提升5-9%,语义对齐提升13-15%,时间同步提升21-52%,舞蹈视频节拍对齐提升28%。大规模主观听感测试也验证了结果一致性。研究表明,基于模态内特征的时间对齐不仅有效,且优于传统成对监督。此外,该方法实现了时间与风格的解耦控制,增强生成可控性。

原文摘要 · Abstract (English)

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned music with disentangled time synchronization and semantic control (e.g., genre, mood) from video while requiring zero video-music pairs at training time. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-ZERO achieves state-of-the-art performance without any paired music-video data, surpassing the strongest prior baselines per metric with 5-9% higher audio quality, 13-15% better semantic alignment, 21-52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Our results validate that temporal alignment through within-modality features is not only effective for video-to-music generation but also leads to better performance than paired cross-modal supervision. Furthermore, our approach enables independent controls for timing and music style (e.g., genre, mood) for more controllable generation.

视频配乐时间对齐零样本生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。