统一建模上下装的扩散模型,实现高保真虚拟试穿。
MuGa-VTON: Multi-Garment Virtual Try-On via Diffusion Transformers with Prompt Customization
- 用共享潜空间联合建模人物与上下装特征
- 在VITON-HD和DressCode上超越现有方法
- 支持提示词定制,可精细调整服装细节
虚拟试穿旨在生成人物穿戴目标服装的逼真图像,需同时保持个人身份特征与服装细节以适用于时尚零售与个性化场景。然而现有方法通常分开展示上下装,依赖复杂预处理,且难以保留纹身、配饰、体型等个体特征,导致真实感与灵活性受限。为此,我们提出MuGa-VTON,一种统一的多服装扩散框架,将上下装与人物身份共同建模于共享潜空间中。具体设计三个核心模块:服装表征模块(GRM)捕捉服装语义,人物表征模块(PRM)编码身份与姿态线索,A-DiT融合模块通过扩散变压器整合服装、人物与文本提示特征。该架构支持基于提示的定制化,仅需少量用户输入即可实现精细服装修改。在VITON-HD与DressCode基准上的大量实验表明,MuGa-VTON在定性与定量评估中均优于现有方法,生成结果具有高保真度与身份一致性,适用于实际虚拟试穿应用。
原文摘要 · Abstract (English)
Virtual try-on seeks to generate photorealistic images of individuals in desired garments, a task that must simultaneously preserve personal identity and garment fidelity for practical use in fashion retail and personalization. However, existing methods typically handle upper and lower garments separately, rely on heavy preprocessing, and often fail to preserve person-specific cues such as tattoos, accessories, and body shape-resulting in limited realism and flexibility. To this end, we introduce MuGa-VTON, a unified multi-garment diffusion framework that jointly models upper and lower garments together with person identity in a shared latent space. Specifically, we proposed three key modules: the Garment Representation Module (GRM) for capturing both garment semantics, the Person Representation Module (PRM) for encoding identity and pose cues, and the A-DiT fusion module, which integrates garment, person, and text-prompt features through a diffusion transformer. This architecture supports prompt-based customization, allowing fine-grained garment modifications with minimal user input. Extensive experiments on the VITON-HD and DressCode benchmarks demonstrate that MuGa-VTON outperforms existing methods in both qualitative and quantitative evaluations, producing high-fidelity, identity-preserving results suitable for real-world virtual try-on applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。