用扩散模型从穿衣照片中重建高保真服装图像
TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models
- 基于SigLIP视觉条件的扩散模型,精准还原服装形状纹理
- 在VITON-HD和Dress Code上优于传统姿态迁移与VTON方法
- 提出新评估指标DISTS,更适合衡量生成质量
本文提出虚拟试穿新任务VTOFF,旨在从单张着装人像中生成标准化服装图像。与传统虚拟试衣(VTON)不同,VTOFF需精确还原服装的形状、纹理和复杂图案,以提升生成模型的评估可靠性。我们提出TryOffDiff,通过引入SigLIP视觉条件改进Stable Diffusion,实现高保真重建。在VITON-HD和Dress Code数据集上的实验表明,TryOffDiff优于适配的姿态迁移与VTON基线。研究发现,传统指标如SSIM难以反映真实重建质量,因此采用DISTS作为可靠评估标准。结果表明,VTOFF可改善电商商品图像,推动生成模型评估发展,并为高保真重建研究提供方向。演示、代码与模型已开源。
原文摘要 · Abstract (English)
This paper introduces Virtual Try-Off (VTOFF), a novel task generating standardized garment images from single photos of clothed individuals. Unlike Virtual Try-On (VTON), which digitally dresses models, VTOFF extracts canonical garment images, demanding precise reconstruction of shape, texture, and complex patterns, enabling robust evaluation of generative model fidelity. We propose TryOffDiff, adapting Stable Diffusion with SigLIP-based visual conditioning to deliver high-fidelity reconstructions. Experiments on VITON-HD and Dress Code datasets show that TryOffDiff outperforms adapted pose transfer and VTON baselines. We observe that traditional metrics such as SSIM inadequately reflect reconstruction quality, prompting our use of DISTS for reliable assessment. Our findings highlight VTOFF's potential to improve e-commerce product imagery, advance generative model evaluation, and guide future research on high-fidelity reconstruction. Demo, code, and models are available at: https://rizavelioglu.github.io/tryoffdiff
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。