arXiv:2601.15872cs.SDcs.CV2026-01

无需人体姿态,直接从舞蹈视频生成匹配音乐,效果领先。

PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation

  • 用舞蹈视频视觉特征替代人体姿态,实现无姿态泛化
  • 在多舞者和非人类舞者场景下仍保持高对齐度与音乐质量
  • 适合跨场景舞蹈音乐生成,尤其适用于复杂真实场景

舞蹈到音乐生成旨在生成与舞蹈动作同步的音乐。现有方法通常依赖单个舞者的人体运动特征和有限的舞蹈-音乐数据集,限制了其在包含多个舞者或非人类舞者的现实场景中的表现与应用。本文提出PF-D2M,一种基于扩散模型的通用舞蹈-音乐生成方法,直接从舞蹈视频中提取视觉特征。该模型采用渐进式训练策略,有效缓解数据稀缺与泛化难题。客观与主观评估均表明,PF-D2M在舞蹈-音乐对齐度与音乐质量方面达到当前最优水平。

原文摘要 · Abstract (English)

Dance-to-music generation aims to generate music that is aligned with dance movements. Existing approaches typically rely on body motion features extracted from a single human dancer and limited dance-to-music datasets, which restrict their performance and applicability to real-world scenarios involving multiple dancers and non-human dancers. In this paper, we propose PF-D2M, a universal diffusion-based dance-to-music generation model that incorporates visual features extracted from dance videos. PF-D2M is trained with a progressive training strategy that effectively addresses data scarcity and generalization challenges. Both objective and subjective evaluations show that PF-D2M achieves state-of-the-art performance in dance-music alignment and music quality.

舞蹈生成扩散模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。