让用户穿想穿的衣服跳舞,5秒高清视频一键生成。
Dress&Dance: Dress up and Dance as You Like It - Technical Preview
- 用注意力机制统一文本、图像、视频输入,提升试穿匹配度。
- 单图输入支持上下装同试,生成1152x720分辨率5秒24帧视频。
- 融合图像与视频数据训练,适合虚拟试衣与内容创作场景。
我们提出 Dress&Dance,一个视频扩散框架,可生成高质量的5秒长、24帧每秒、1152x720分辨率的虚拟试穿视频,让用户在参考视频动作下穿着指定衣物动态展示。该方法仅需一张用户图像,支持上衣、下装及连体服,且可在单次推理中同时完成上下装试穿。核心是 CondNet——一种新型条件网络,通过注意力机制融合多模态输入(文本、图像、视频),显著提升衣物对齐与动作保真度。CondNet 在异构数据上分阶段渐进训练,结合少量视频数据与大量图像数据。Dress&Dance 超越现有开源与商业方案,实现高画质且灵活的虚拟试穿体验。
原文摘要 · Abstract (English)
We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments while moving in accordance with a given reference video. Our approach requires a single user image and supports a range of tops, bottoms, and one-piece garments, as well as simultaneous tops and bottoms try-on in a single pass. Key to our framework is CondNet, a novel conditioning network that leverages attention to unify multi-modal inputs (text, images, and videos), thereby enhancing garment registration and motion fidelity. CondNet is trained on heterogeneous training data, combining limited video data and a larger, more readily available image dataset, in a multistage progressive manner. Dress&Dance outperforms existing open source and commercial solutions and enables a high quality and flexible try-on experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。