让音乐精准匹配视频节奏与情绪,实现画面转场与鼓点严丝合缝。
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
- 分层解析视频信息,用跨模态注意力融合语义与时间线索
- 帧级动态对齐视觉切换与音乐节拍,实现精确到毫秒的节奏同步
- 专为电商广告设计新数据集,适合做视频配乐生成的研究者
视频配乐旨在生成能增强视听沉浸感的背景音乐。现有方法存在两大缺陷:一是视频细节表征不全,导致音画对齐弱;二是缺乏精确的时间与节奏对应,尤其难以实现节拍同步。为此,我们提出视频音乐对齐模型VeM,一种潜空间音乐扩散模型,可生成在语义、时间与节奏上均与输入视频高度对齐的高质量原声。为全面捕捉视频细节,VeM采用分层视频解析机制,作为音乐指挥官,在多层级跨模态信息间进行协调。特定模态编码器结合故事板引导的交叉注意力(SG-CAtt),通过位置与持续时间编码保持时间连贯性并整合语义线索。为提升节奏精度,引入帧级过渡-节拍对齐器与适配器(TB-As),动态将视觉场景切换与音乐节拍对齐。我们还构建了一个来自电商广告与视频分享平台的新视频-音乐配对数据集,要求更高的转场-节拍同步标准,并设计了针对性评估指标。实验表明,该方法在语义相关性与节奏精确性方面显著优于现有方法。
原文摘要 · Abstract (English)
Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。