arXiv:2501.16757cs.CV2025-01中稿 · PRCV 2025被引 4

用单一Transformer模型实现高保真虚拟试穿,省去冗余网络

ITVTON: Virtual Try-On Diffusion Transformer Based on Integrated Image and Text

  • 仅用一个Diffusion Transformer生成器,融合图像与文本信息
  • 在IGPair数据集上10,257对图像测试中表现优异,生成更真实细节
  • 只训练单个DiT块的注意力参数,显著降低计算开销

虚拟试穿旨在无缝将衣物贴合到人物图像上,近期基于扩散模型的方法取得显著进展。然而,现有方法通常依赖重复的主干网络或额外的图像编码器提取衣物特征,导致计算开销和网络复杂度增加。本文提出ITVTON,一种基于扩散Transformer(DiT)的高效框架,仅使用一个生成器提升图像保真度。通过沿宽度方向拼接衣物与人物图像,并融合两者文本描述,ITVTON有效捕捉衣物-人物交互关系,同时保持真实感。为进一步降低计算成本,训练仅限于单个扩散Transformer(Single-DiT)块内的注意力参数。大量实验表明,ITVTON在定性与定量指标上均优于基线方法,树立了虚拟试穿新标准。此外,在包含10,257对图像的IGPair数据集上的实验验证了其在真实场景中的鲁棒性。

原文摘要 · Abstract (English)

Virtual try-on, which aims to seamlessly fit garments onto person images, has recently seen significant progress with diffusion-based models. However, existing methods commonly resort to duplicated backbones or additional image encoders to extract garment features, which increases computational overhead and network complexity. In this paper, we propose ITVTON, an efficient framework that leverages the Diffusion Transformer (DiT) as its single generator to improve image fidelity. By concatenating garment and person images along the width dimension and incorporating textual descriptions from both, ITVTON effectively captures garment-person interactions while preserving realism. To further reduce computational cost, we restrict training to the attention parameters within a single Diffusion Transformer (Single-DiT) block. Extensive experiments demonstrate that ITVTON surpasses baseline methods both qualitatively and quantitatively, setting a new standard for virtual try-on. Moreover, experiments on 10,257 image pairs from IGPair confirm its robustness in real-world scenarios.

虚拟试穿扩散模型Transformer图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。