让视频人物零样本唱歌跳舞,还能区分主唱和听众角色
SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning

- 用角色感知音频条件控制声乐动作与舞蹈
- 实现音画对齐强、口型同步好、参数更少的生成效果
- 适合做个性化音乐视频创作或跨模态生成研究
从参考图像、文本提示和音频轨道生成个性化歌舞视频,需实现音乐驱动的身体动作。歌唱舞蹈还需视觉主体同步发声。现有方法多关注编舞,语音驱动模型通常假设主体发出输入语音,该联合场景仍待探索。我们提出SingDance,一个统一的视频扩散框架,将可控声乐动作定义为语义角色:主体可为发声者(源)或听者(监听者)。硬压缩路由选择任务相关的语音、音乐和角色条件,通过帧级联合音频注入进行组合;源与听者共享同一语音路径。训练采用非对称监督:屏幕内说话视频与精心筛选的离屏对话响应视频建立角色控制,而纯乐器与歌曲舞蹈视频建立音乐驱动身体动作。目标‘歌曲+源’配置在训练中从未出现。推理时,将歌曲设为源角色,即可组合已独立学习的发声与歌曲驱动舞蹈能力,实现组合式零样本歌唱舞蹈生成。实验表明其运动-节拍对齐强、视觉保真度高,能可靠切换声乐动作且保持音乐对齐的身体动作,唇形同步性能优于最强基线,生成参数显著减少。
原文摘要 · Abstract (English)
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion--beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。