让音乐生成更听指令,还能自动校正错误。
VIBE: Video Instruction-aligned Background music gEneration

- 用动态跨层条件机制连接规划与细化模块。
- 支持节奏、调性等硬约束和音乐性等软要求。
- 适合需要精准控制背景音乐的创作者或开发者。
当前视频到音乐(V2M)模型缺乏语义控制能力,且无法惩罚指令违规,主要源于其依赖重建目标以及扩散自回归(DAR)架构中静态跨模态条件的表征瓶颈。为此,我们提出VIBE,一种新颖的文本与视频到音乐(T+V2M)生成模型,其创新包括:(1) 条件连接(Conditioning Connection),一种深度交叉层条件机制,可动态连接规划头与扩散精修头;(2) 综合奖励建模分类法,通过结构化的五阶段训练课程,同时优化硬性可验证约束(如节奏、调性)与软性主观质量(如音乐性、多模态对齐)。在音频-视觉对齐、指令遵循度及音频质量指标上进行评估,并辅以主观人类评测,结果表明VIBE在可控性和指令遵循方面表现更优,同时在生成保真度与多模态对齐上与多数基线模型相当。
原文摘要 · Abstract (English)
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Autoregressive (DAR) architectures. To resolve this, we introduce VIBE, a novel text-and-video-to-music (T+V2M) generation model that leverages: (1) Conditioning Connection, a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and (2) a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints (e.g., tempo, key) and soft, subjective qualities (e.g., musicality, multimodal alignment) with a structured 5-stage training curriculum. Upon evaluation using audio-visual alignment, instruction following, and audio quality metrics, along with a subjective human evaluation study, we observe that VIBE demonstrates enhanced controllability and instruction adherence while performing comparably to most evaluated baselines on generation fidelity and multimodal alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。