arXiv:2411.01593cs.CV2024-11

用海量无配对服装图训练高保真虚拟试穿,突破数据限制。

High-Fidelity Virtual Try-on with Large-Scale Unpaired Learning

  • 通过逆向变形生成服装标准模板,构建伪配对数据。
  • 在1024×768分辨率上优于现有方法,细节还原更真实。
  • 适合电商、时尚应用,对多种穿搭风格泛化性强。

虚拟试穿(VTON)将目标服装图像转移到参考人体上,服装保真度是下游电商业务的关键需求。然而,现有方法在高保真试穿方面仍存在不足,原因在于穿搭风格多样性(如衣物被裤子遮挡或姿态导致变形)与训练数据有限的配对样本之间的矛盾。本文提出一种新框架——增强型虚拟试穿(BVTON),利用大规模无配对学习实现高保真试穿。核心思路是从海量时尚图像中可靠构建伪试穿配对。首先,提出组合式规范流,将模特穿着的服装映射为类似店铺内展示的形态,称为标准代理;每个服装部件(袖子、躯干)被逆向变形为店铺样式的形状,以组合构造标准代理。其次,设计分层掩码生成模块,基于标准代理训练生成精确语义布局,替代传统流程中的店铺服装以提升训练效果。最后,通过随机错位的模特服装构建伪训练对,设计无配对试穿合成器,生成精细的皮肤纹理和衣物边界。在高分辨率(1024×768)数据集上的大量实验表明,该方法在定性和定量上均优于当前最优方法。值得注意的是,BVTON在不同穿搭风格和数据源下展现出优异的泛化性与可扩展性。

原文摘要 · Abstract (English)

Virtual try-on (VTON) transfers a target clothing image to a reference person, where clothing fidelity is a key requirement for downstream e-commerce applications. However, existing VTON methods still fall short in high-fidelity try-on due to the conflict between the high diversity of dressing styles (\eg clothes occluded by pants or distorted by posture) and the limited paired data for training. In this work, we propose a novel framework \textbf{Boosted Virtual Try-on (BVTON)} to leverage the large-scale unpaired learning for high-fidelity try-on. Our key insight is that pseudo try-on pairs can be reliably constructed from vastly available fashion images. Specifically, \textbf{1)} we first propose a compositional canonicalizing flow that maps on-model clothes into pseudo in-shop clothes, dubbed canonical proxy. Each clothing part (sleeves, torso) is reversely deformed into an in-shop-like shape to compositionally construct the canonical proxy. \textbf{2)} Next, we design a layered mask generation module that generates accurate semantic layout by training on canonical proxy. We replace the in-shop clothes used in conventional pipelines with the derived canonical proxy to boost the training process. \textbf{3)} Finally, we propose an unpaired try-on synthesizer by constructing pseudo training pairs with randomly misaligned on-model clothes, where intricate skin texture and clothes boundaries can be generated. Extensive experiments on high-resolution ($1024\times768$) datasets demonstrate the superiority of our approach over state-of-the-art methods both qualitatively and quantitatively. Notably, BVTON shows great generalizability and scalability to various dressing styles and data sources.

虚拟试穿无配对学习高保真服饰生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。