提出Dual-UNet扩散模型,解决虚拟试衣中服装还原难题。
What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction
- 采用Dual-UNet架构,融合多尺度特征与条件控制
- 在VITON-HD和DressCode上实现DISTS下降9.5%的最优性能
- 为服装重建提供可复现基线,适合生成式时尚研究者
虚拟试衣(VTON)发展迅速,但其逆问题——从穿在身上的图像重建原始服装(即虚拟试衣还原,VTOFF)仍不明确。本文通过分析基于扩散模型的VTON与通用潜空间扩散模型(LDMs)策略,聚焦于Dual-UNet扩散模型架构,考察三方面设计:(i) 生成主干网络(比较Stable Diffusion变体);(ii) 条件输入(对比不同掩码设计、有无掩码输入及高层语义特征);(iii) 损失函数与训练策略(评估注意力辅助损失、感知目标及分阶段课程调度)。大量实验揭示各配置间的权衡。在VITON-HD与DressCode数据集上,本框架达当前最优表现,主指标DISTS降低9.5%,在LPIPS、FID、KID、SSIM上也具竞争力,既提供了更强基线,也为未来虚拟试衣研究提供指导。
原文摘要 · Abstract (English)
Virtual Try-On (VTON) has seen rapid advancements, providing a strong foundation for generative fashion tasks. However, the inverse problem, Virtual Try-Off (VTOFF)-aimed at reconstructing the canonical garment from a draped-on image-remains a less understood domain, distinct from the heavily researched field of VTON. In this work, we seek to establish a robust architectural foundation for VTOFF by studying and adapting various diffusion-based strategies from VTON and general Latent Diffusion Models (LDMs). We focus our investigation on the Dual-UNet Diffusion Model architecture and analyze three axes of design: (i) Generation Backbone: comparing Stable Diffusion variants; (ii) Conditioning: ablating different mask designs, masked/unmasked inputs for image conditioning, and the utility of high-level semantic features; and (iii) Losses and Training Strategies: evaluating the impact of the auxiliary attention-based loss, perceptual objectives and multi-stage curriculum schedules. Extensive experiments reveal trade-offs across various configuration options. Evaluated on VITON-HD and DressCode datasets, our framework achieves state-of-the-art performance with a drop of 9.5\% on the primary metric DISTS and competitive performance on LPIPS, FID, KID, and SSIM, providing both stronger baselines and insights to guide future Virtual Try-Off research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。