分两阶段实现高保真虚拟试穿,精准对齐衣物与人体姿态。
DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On
- 先变形对齐衣物,再细化纹理细节,解耦生成流程。
- 在多个数据集上超越现有方法,保留褶皱、纹理等关键细节。
- 适合电商、数字时装领域,对复杂姿态和款式泛化性强。
虚拟试穿(VTON)旨在合成人物穿着目标服饰的逼真图像,在电子商务和数字时尚中应用广泛。尽管潜在扩散模型的进展显著提升了视觉质量,现有方法仍难以保持细粒度服饰细节、实现精确的服饰-人体对齐、维持推理效率,并泛化至多样姿态与服装风格。为此,我们提出DiffFit,一种新颖的两阶段潜在扩散框架,用于高保真虚拟试穿。DiffFit采用渐进生成策略:第一阶段进行几何感知的衣物形变,通过细粒度变形与姿态适配,将衣物对齐至目标人体;第二阶段通过跨模态条件扩散模型,融合形变后的衣物、原始衣物外观及目标人物图像,实现高质量纹理精细化。通过解耦几何对齐与外观精修,DiffFit有效降低任务复杂度,提升生成稳定性和视觉真实感。其在保留服饰特有属性如纹理、褶皱、光照方面表现优异,同时确保与人体的精确对齐。大规模VTON基准测试表明,DiffFit在定量指标与主观评价上均优于现有最先进方法。
原文摘要 · Abstract (English)
Virtual try-on (VTON) aims to synthesize realistic images of a person wearing a target garment, with broad applications in e-commerce and digital fashion. While recent advances in latent diffusion models have substantially improved visual quality, existing approaches still struggle with preserving fine-grained garment details, achieving precise garment-body alignment, maintaining inference efficiency, and generalizing to diverse poses and clothing styles. To address these challenges, we propose DiffFit, a novel two-stage latent diffusion framework for high-fidelity virtual try-on. DiffFit adopts a progressive generation strategy: the first stage performs geometry-aware garment warping, aligning the garment with the target body through fine-grained deformation and pose adaptation. The second stage refines texture fidelity via a cross-modal conditional diffusion model that integrates the warped garment, the original garment appearance, and the target person image for high-quality rendering. By decoupling geometric alignment and appearance refinement, DiffFit effectively reduces task complexity and enhances both generation stability and visual realism. It excels in preserving garment-specific attributes such as textures, wrinkles, and lighting, while ensuring accurate alignment with the human body. Extensive experiments on large-scale VTON benchmarks demonstrate that DiffFit achieves superior performance over existing state-of-the-art methods in both quantitative metrics and perceptual evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。