用正负节奏信息提升舞步到音乐生成的同步性与质量
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
- 引入正负条件扩散模型,双向利用舞步节奏线索
- 在AIST++和TikTok数据集上实现更精准的节拍对齐
- 适合音乐生成与跨模态对齐研究者参考
条件扩散模型在跨模态合成中表现优异,通过时间条件U-Net结合交叉注意力机制,可实现条件输入与生成输出间的强对齐。本文聚焦于根据给定舞蹈视频生成同步音乐的问题。考虑到双向引导更有利于扩散模型训练,我们提出PN-Diffusion,同时以正向节奏信息和反向节奏信息作为条件,设计双重扩散与逆过程。具体而言,为训练序列多模态U-Net结构,该模型包含正条件下的噪声预测目标和额外的负条件噪声预测目标。为准确提取并选择正负条件,我们巧妙利用舞蹈视频中的时间相关性,分别通过正向播放和反向播放捕捉正负节奏线索。在AIST++和TikTok舞蹈视频数据集上,通过主观与客观评估输入输出间舞蹈-音乐节拍对齐程度及生成音乐质量,实验结果表明,本模型优于当前最先进舞蹈到音乐生成方法。
原文摘要 · Abstract (English)
Conditional diffusion models have gained increasing attention since their impressive results for cross-modal synthesis, where the strong alignment between conditioning input and generated output can be achieved by training a time-conditioned U-Net augmented with cross-attention mechanism. In this paper, we focus on the problem of generating music synchronized with rhythmic visual cues of the given dance video. Considering that bi-directional guidance is more beneficial for training a diffusion model, we propose to enhance the quality of generated music and its synchronization with dance videos by adopting both positive rhythmic information and negative ones (PN-Diffusion) as conditions, where a dual diffusion and reverse processes is devised. Specifically, to train a sequential multi-modal U-Net structure, PN-Diffusion consists of a noise prediction objective for positive conditioning and an additional noise prediction objective for negative conditioning. To accurately define and select both positive and negative conditioning, we ingeniously utilize temporal correlations in dance videos, capturing positive and negative rhythmic cues by playing them forward and backward, respectively. Through subjective and objective evaluations of input-output correspondence in terms of dance-music beat alignment and the quality of generated music, experimental results on the AIST++ and TikTok dance video datasets demonstrate that our model outperforms SOTA dance-to-music generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。