arXiv:2604.17656cs.SDcs.AI2026-04被引 2

用文本控制视频配乐,生成更快更准。

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

论文配图:Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
图 1 · 摘自论文原文
  • 先用文本+视频规划音乐结构,再用扩散模型细化音质。
  • 比现有方法快2.21倍,跨域生成效果更优。
  • 适合需要精细控制配乐风格的创作者使用。

视频到音乐(V2M)是为输入视频生成背景音乐的基础任务。现有V2M模型通常仅依赖视觉条件,缺乏语义和风格上的用户可控性。本文提出Video-Robin,一种新型文本条件下的视频到音乐生成模型,可实现快速、高质量且语义对齐的音乐生成。为平衡音乐保真度与语义理解,Video-Robin结合自回归规划与基于扩散的合成:自回归模块通过语义对齐视频与文本输入,生成高层音乐潜在表示;这些表示随后由局部扩散变换器精炼为连贯高保真的音乐。通过将语义驱动的规划解耦至扩散合成中,Video-Robin在不牺牲音频真实性的前提下,实现细粒度创作控制。该模型在分布内与分布外基准上均超越仅接受视频输入或额外特征条件的基线模型,推理速度达当前最优水平的2.21倍。论文接受后将开源全部代码与模型。

原文摘要 · Abstract (English)

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic and stylistic controllability to the end user. In this paper, we present Video-Robin, a novel text-conditioned video-to-music generation model that enables fast, high-quality, semantically aligned music generation for video content. To balance musical fidelity and semantic understanding, Video-Robin integrates autoregressive planning with diffusion-based synthesis. Specifically, an autoregressive module models global structure by semantically aligning visual and textual inputs to produce high-level music latents. These latents are subsequently refined into coherent, high-fidelity music using local Diffusion Transformers. By factoring semantically driven planning into diffusion-based synthesis, Video-Robin enables fine-grained creator control without sacrificing audio realism. Our proposed model outperforms baselines that solely accept video input and additional feature conditioned baselines on both in-distribution and out-of-distribution benchmarks with a 2.21x speed in inference compared to SOTA. We will open-source everything upon paper acceptance.

视频配乐扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。