用扩散变压器统一实现试穿与试所有,支持多种商品和精细编辑。
DiT-VTON: Diffusion Transformer Framework for Unified Multi-Category Virtual Try-On and Virtual Try-All with Integrated Image Editing
- 基于扩散Transformer架构,融合多模态条件输入提升图像生成质量。
- 在数千类商品上超越现有方法,细节保留更佳且无需额外编码器。
- 支持姿态保持、局部修改、纹理迁移等高级编辑功能,适合电商应用。
电子商务的快速发展推动了虚拟试穿(VTO)技术的需求,使用户能真实预览服装叠加在自身图像上的效果。尽管近期取得进展,现有VTO模型仍面临细粒度细节保留差、对真实图像鲁棒性不足、采样效率低、图像编辑能力弱以及跨品类泛化能力差等问题。本文提出DiT-VTON,一种基于扩散Transformer(DiT)的新颖VTO框架,将擅长文本条件图像生成的DiT adapted为图像条件任务。系统评估了多种配置,包括上下文标记拼接、通道拼接及ControlNet集成,以确定最优方案。为增强鲁棒性,模型在扩展数据集上训练,涵盖多样背景、非结构化参考图和非服装类别,验证了数据规模对VTO适应性的益处。此外,DiT-VTON将任务拓展至虚拟试所有(VTA),可处理广泛产品类别,并支持姿态保持、局部编辑、纹理转移和对象级定制等高级图像编辑功能。实验表明,该模型在VITON-HD上优于现有最先进方法,细节保留更优且不依赖额外条件编码器;同时在跨越数千个产品类别的多样化数据集上,也优于具备VTA与编辑能力的模型。
原文摘要 · Abstract (English)
The rapid growth of e-commerce has intensified the demand for Virtual Try-On (VTO) technologies, enabling customers to realistically visualize products overlaid on their own images. Despite recent advances, existing VTO models face challenges with fine-grained detail preservation, robustness to real-world imagery, efficient sampling, image editing capabilities, and generalization across diverse product categories. In this paper, we present DiT-VTON, a novel VTO framework that leverages a Diffusion Transformer (DiT), renowned for its performance on text-conditioned image generation, adapted here for the image-conditioned VTO task. We systematically explore multiple DiT configurations, including in-context token concatenation, channel concatenation, and ControlNet integration, to determine the best setup for VTO image conditioning. To enhance robustness, we train the model on an expanded dataset encompassing varied backgrounds, unstructured references, and non-garment categories, demonstrating the benefits of data scaling for VTO adaptability. DiT-VTON also redefines the VTO task beyond garment try-on, offering a versatile Virtual Try-All (VTA) solution capable of handling a wide range of product categories and supporting advanced image editing functionalities such as pose preservation, localized editing, texture transfer, and object-level customization. Experimental results show that our model surpasses state-of-the-art methods on VITON-HD, achieving superior detail preservation and robustness without reliance on additional condition encoders. It also outperforms models with VTA and image editing capabilities on a diverse dataset spanning thousands of product categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。