arXiv:2501.06769cs.CV2025-01被引 1

用姿态引导扩散模型,实现动态姿势下逼真虚拟试穿。

ODPG: Outfitting Diffusion with Pose Guided Condition

  • 多条件输入融合姿态、服装与外观特征,端到端生成
  • 在FashionTryOn和DeepFashion子集上生成细节丰富的真实图像
  • 无需显式服装变形,适合电商与数字时尚应用

虚拟试穿(VTON)技术让用户在不实际试穿的情况下预览服装效果,随着数字化和在线购物的发展而受到关注。传统VTON方法通常使用生成对抗网络(GAN)和扩散模型,在实现高真实感和处理动态姿势方面面临挑战。本文提出一种新方法——姿态引导的扩散模型(ODPG),该方法在去噪过程中引入多个条件输入,利用潜空间扩散模型,将服装、姿态和外观图像转换为潜在特征,并在基于UNet的去噪模型中融合这些特征,实现对动态姿势人体图像上服装的非显式合成。在FashionTryOn和DeepFashion数据集子集上的实验表明,ODPG能生成具有精细纹理细节的逼真虚拟试穿图像,采用端到端架构,无需显式服装变形过程。未来工作将聚焦于生成视频形式的VTON输出,并将本文提出的注意力机制应用于数据有限的其他领域。

原文摘要 · Abstract (English)

Virtual Try-On (VTON) technology allows users to visualize how clothes would look on them without physically trying them on, gaining traction with the rise of digitalization and online shopping. Traditional VTON methods, often using Generative Adversarial Networks (GANs) and Diffusion models, face challenges in achieving high realism and handling dynamic poses. This paper introduces Outfitting Diffusion with Pose Guided Condition (ODPG), a novel approach that leverages a latent diffusion model with multiple conditioning inputs during the denoising process. By transforming garment, pose, and appearance images into latent features and integrating these features in a UNet-based denoising model, ODPG achieves non-explicit synthesis of garments on dynamically posed human images. Our experiments on the FashionTryOn and a subset of the DeepFashion dataset demonstrate that ODPG generates realistic VTON images with fine-grained texture details across various poses, utilizing an end-to-end architecture without the need for explicit garment warping processes. Future work will focus on generating VTON outputs in video format and on applying our attention mechanism, as detailed in the Method section, to other domains with limited data.

虚拟试穿扩散模型姿态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。