用音乐和文字驱动图片跳舞,生成同步动作的个性化舞蹈视频
Every Image Listens, Every Image Dances: Music-Driven Image Animation
- 输入音乐+文字,直接让图片角色随音乐起舞
- 生成2904段同步音乐节奏的舞蹈视频,动作连贯自然
- 无需姿态或深度数据,普通人也能轻松创作
图像动画已成为多模态研究的热门方向,重点在于从参考图像生成视频。尽管先前工作多聚焦于文本引导的通用视频生成,音乐驱动的舞蹈视频生成仍鲜有探索。本文提出MuseDance,一种端到端模型,利用音乐与文本双重输入,实现参考图像的动态化。该方法可生成符合文本描述、且角色动作与音乐节奏同步的个性化视频。相比现有方法,MuseDance无需复杂运动引导输入(如姿态序列或深度图),使灵活创意视频生成对各类用户均易用。为推动该领域发展,我们构建了一个新的多模态数据集,包含2,904个带有背景音乐和文本描述的舞蹈视频。模型采用基于扩散的方法,实现强泛化性、精确控制与时间一致性,为音乐驱动图像动画任务设立了新基准。
原文摘要 · Abstract (English)
Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven dance video generation remains underexplored. In this paper, we introduce MuseDance, an innovative end-to-end model that animates reference images using both music and text inputs. This dual input enables MuseDance to generate personalized videos that follow text descriptions and synchronize character movements with the music. Unlike existing approaches, MuseDance eliminates the need for complex motion guidance inputs, such as pose or depth sequences, making flexible and creative video generation accessible to users of all expertise levels. To advance research in this field, we present a new multimodal dataset comprising 2,904 dance videos with corresponding background music and text descriptions. Our approach leverages diffusion-based methods to achieve robust generalization, precise control, and temporal consistency, setting a new baseline for the music-driven image animation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。