让舞蹈生成音乐更精准,解决节奏错位和对齐问题
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
- 用多尺度小波分析提取细粒度节奏特征,自适应调整关节权重
- 引入可学习上下文查询,解决特征下采样导致的时间错位
- 在AIST++和TikTok数据集上效果优于现有方法
舞蹈生成音乐(D2M)旨在自动生成与舞蹈动作在节奏和时间上对齐的音乐。现有方法通常依赖粗粒度节奏嵌入,如全局运动特征或二值化关节节奏值,丢失了细粒度运动线索,导致节奏对齐较弱。此外,特征下采样引入的时间错位进一步阻碍了舞蹈与音乐间的精确同步。为此,我们提出基于扩散变换器的GACA-DiT框架,包含两个新模块:首先,风格自适应节奏提取模块结合多尺度时序小波分析与空间相位直方图,通过自适应关节加权捕捉细粒度、风格相关的节奏模式;其次,上下文感知时间对齐模块利用可学习上下文查询,将音乐潜在表示与相关舞蹈节奏特征对齐。在AIST++和TikTok数据集上的大量实验表明,GACA-DiT在客观指标和人类评估中均优于当前最先进方法。
原文摘要 · Abstract (English)
Dance-to-music (D2M) generation aims to automatically compose music that is rhythmically and temporally aligned with dance movements. Existing methods typically rely on coarse rhythm embeddings, such as global motion features or binarized joint-based rhythm values, which discard fine-grained motion cues and result in weak rhythmic alignment. Moreover, temporal mismatches introduced by feature downsampling further hinder precise synchronization between dance and music. To address these problems, we propose \textbf{GACA-DiT}, a diffusion transformer-based framework with two novel modules for rhythmically consistent and temporally aligned music generation. First, a \textbf{genre-adaptive rhythm extraction} module combines multi-scale temporal wavelet analysis and spatial phase histograms with adaptive joint weighting to capture fine-grained, genre-specific rhythm patterns. Second, a \textbf{context-aware temporal alignment} module resolves temporal mismatches using learnable context queries to align music latents with relevant dance rhythm features. Extensive experiments on the AIST++ and TikTok datasets demonstrate that GACA-DiT outperforms state-of-the-art methods in both objective metrics and human evaluation. Project page: https://beria-moon.github.io/GACA-DiT/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。