用扩散模型生成逼真单人舞蹈视频,提升细节与动作连贯性。
DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model
- 通过外观增强和姿态渲染模块,融合参考图与动作信息。
- 在TikTok-3K数据集上训练,生成视频质量优于现有方法。
- 适合关注舞蹈生成、视频合成的开发者与研究者。
扩散模型在视觉生成任务中展现出强大能力,尤其在动态视频生成方面。然而,视频生成对质量要求更高,且需保证连续性。涉及人体动作的视频(如舞蹈)因人体运动自由度高,生成难度更大。本文提出DANCER框架,基于最新的Stable Video Diffusion模型,实现单人舞蹈的逼真合成。针对视频生成通常依赖参考图像和视频序列的特点,设计了外观增强模块(AEM)以强化参考图细节,并引入姿态渲染模块(PRM)从额外域中提取姿态条件,扩展运动引导能力。为提升模型性能,还从互联网收集大量视频数据,构建新数据集TikTok-3K用于训练。在真实数据集上的大量实验表明,该模型性能超越当前最优方法。所有数据与代码将在录用后公开。
原文摘要 · Abstract (English)
Recently, diffusion models have shown their impressive ability in visual generation tasks. Besides static images, more and more research attentions have been drawn to the generation of realistic videos. The video generation not only has a higher requirement for the quality, but also brings a challenge in ensuring the video continuity. Among all the video generation tasks, human-involved contents, such as human dancing, are even more difficult to generate due to the high degrees of freedom associated with human motions. In this paper, we propose a novel framework, named as DANCER (Dance ANimation via Condition Enhancement and Rendering with Diffusion Model), for realistic single-person dance synthesis based on the most recent stable video diffusion model. As the video generation is generally guided by a reference image and a video sequence, we introduce two important modules into our framework to fully benefit from the two inputs. More specifically, we design an Appearance Enhancement Module (AEM) to focus more on the details of the reference image during the generation, and extend the motion guidance through a Pose Rendering Module (PRM) to capture pose conditions from extra domains. To further improve the generation capability of our model, we also collect a large amount of video data from Internet, and generate a novel datasetTikTok-3K to enhance the model training. The effectiveness of the proposed model has been evaluated through extensive experiments on real-world datasets, where the performance of our model is superior to that of the state-of-the-art methods. All the data and codes will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。