用扩散模型生成高保真试穿视频,保持人物身份与动作连贯。
Fashion-VDM: Video Diffusion Model for Virtual Try-On
- 基于扩散模型设计新架构,分步控制服装与人体输入。
- 单次生成64帧512像素视频,提升细节与时间一致性。
- 数据少时联合训练仍有效,适合电商试穿场景。
我们提出Fashion-VDM,一种用于生成虚拟试穿视频的视频扩散模型(VDM)。给定一件服装图像和一段人物视频,该方法旨在生成高质量的试穿视频,使人物穿着指定服装的同时保持其身份和动作特征。虽然基于图像的虚拟试穿已取得显著成果,但现有视频虚拟试穿(VVT)方法在服装细节和时间一致性方面仍有不足。为此,我们提出一种基于扩散的视频试穿架构,采用分离式无分类器引导以增强条件输入控制,并引入渐进式时间训练策略,实现单次生成64帧、512像素的视频。我们还验证了图像-视频联合训练在视频数据有限时的有效性。定性和定量实验表明,本方法在视频虚拟试穿任务上达到新基准。更多结果请访问项目页:https://johannakarras.github.io/Fashion-VDM。
原文摘要 · Abstract (English)
We present Fashion-VDM, a video diffusion model (VDM) for generating virtual try-on videos. Given an input garment image and person video, our method aims to generate a high-quality try-on video of the person wearing the given garment, while preserving the person's identity and motion. Image-based virtual try-on has shown impressive results; however, existing video virtual try-on (VVT) methods are still lacking garment details and temporal consistency. To address these issues, we propose a diffusion-based architecture for video virtual try-on, split classifier-free guidance for increased control over the conditioning inputs, and a progressive temporal training strategy for single-pass 64-frame, 512px video generation. We also demonstrate the effectiveness of joint image-video training for video try-on, especially when video data is limited. Our qualitative and quantitative experiments show that our approach sets the new state-of-the-art for video virtual try-on. For additional results, visit our project page: https://johannakarras.github.io/Fashion-VDM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。