让多人在混乱控制信号下跳舞,还能保持各自身份不变。
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
- 用追踪掩码融合姿态热图,每步去噪都绑定身份与动作
- 在26小时双人滑冰数据上实现7000+身份精准保留
- 适合数字人表演、人机交互仿真等需要多角色互动的场景
可控视频生成虽快速进展,但当多个演员需在噪声控制信号下移动、互动并交换位置时,现有系统仍表现不佳。我们提出DanceTogether,首个端到端扩散框架,仅需一张参考图和独立的姿态掩码流,即可生成长时序、逼真的多人互动视频,并严格保持每个角色的身份特征。创新的MaskPoseAdapter在每一步去噪中融合鲁棒的追踪掩码与语义丰富但噪声较大的姿态热图,消除身份漂移和外观渗漏问题。为支持大规模训练与评估,我们构建了三个新资源:(i) PairFS-4K,包含26小时双人滑冰视频,涵盖7000多个不同身份;(ii) HumanRob-300,一小时人形机器人交互数据集,用于跨域快速迁移;(iii) TogetherVideoBench,一个以DanceTogEval-100测试集为核心的三轨基准,覆盖舞蹈、拳击、摔跤、瑜伽和花样滑冰。在TogetherVideoBench上,DanceTogether显著超越现有方法。此外,仅需一小时微调即可生成可信的人机交互视频,证明其在具身人工智能与人机交互任务中的强泛化能力。大量消融实验确认,持续的身份-动作绑定是性能提升的关键。综上,我们的模型、数据集与基准将可控视频生成从单人编排推进至可组合控制的多角色互动阶段,为数字制作、仿真及具身智能开辟新路径。视频演示与代码已公开于 https://DanceTog.github.io/。
原文摘要 · Abstract (English)
Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image plus independent pose-mask streams into long, photorealistic videos while strictly preserving every identity. A novel MaskPoseAdapter binds "who" and "how" at every denoising step by fusing robust tracking masks with semantically rich-but noisy-pose heat-maps, eliminating the identity drift and appearance bleeding that plague frame-wise pipelines. To train and evaluate at scale, we introduce (i) PairFS-4K, 26 hours of dual-skater footage with 7,000+ distinct IDs, (ii) HumanRob-300, a one-hour humanoid-robot interaction set for rapid cross-domain transfer, and (iii) TogetherVideoBench, a three-track benchmark centered on the DanceTogEval-100 test suite covering dance, boxing, wrestling, yoga, and figure skating. On TogetherVideoBench, DanceTogether outperforms the prior arts by a significant margin. Moreover, we show that a one-hour fine-tune yields convincing human-robot videos, underscoring broad generalization to embodied-AI and HRI tasks. Extensive ablations confirm that persistent identity-action binding is critical to these gains. Together, our model, datasets, and benchmark lift CVG from single-subject choreography to compositionally controllable, multi-actor interaction, opening new avenues for digital production, simulation, and embodied intelligence. Our video demos and code are available at https://DanceTog.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。