用时空图Mamba模型,把音乐转成自然舞蹈视频。
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
- 设计时空图Mamba模块,从音乐生成连贯的骨骼序列。
- 自监督正则化网络提升骨骼到视频的转换质量。
- 构建5.5万条数据集,性能显著优于现有方法。
我们提出一种新型时空图Mamba(STG-Mamba)用于音乐引导的舞蹈视频合成任务,即把输入音乐转换为舞蹈视频。STG-Mamba包含两个映射:音乐到骨骼序列的转换,以及骨骼序列到视频的转换。在音乐到骨骼转换中,引入新颖的时空图Mamba(STGM)模块,有效从音乐构建骨骼序列,捕捉关节在空间和时间维度上的依赖关系。在骨骼到视频转换中,提出一种新的自监督正则化网络,结合条件图像将生成的骨骼序列转化为舞蹈视频。最后,我们从互联网收集了一个新的骨骼到视频转换数据集,包含54,944个视频片段。大量实验表明,STG-Mamba显著优于现有方法。
原文摘要 · Abstract (English)
We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-skeleton translation and skeleton-to-video translation. In the music-to-skeleton translation, we introduce a novel spatial-temporal graph Mamba (STGM) block to effectively construct skeleton sequences from the input music, capturing dependencies between joints in both the spatial and temporal dimensions. For the skeleton-to-video translation, we propose a novel self-supervised regularization network to translate the generated skeletons, along with a conditional image, into a dance video. Lastly, we collect a new skeleton-to-video translation dataset from the Internet, containing 54,944 video clips. Extensive experiments demonstrate that STG-Mamba achieves significantly better results than existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。