提出可叠加衣物的虚拟试穿方法,解决传统方法无法处理多层穿搭的问题。
Layering Virtual Try-On

- 分两阶段训练:先用合成数据学通用试穿规律,再用小规模数据学叠穿逻辑。
- 在新基准上达到最先进效果,且在传统试穿任务上展现零样本能力。
- 适合研究多层服饰生成、时尚应用开发的开发者和研究人员。
现实中穿搭是层层叠加的过程,如外套罩衬衫或反复增减衣物,而现有虚拟试穿(VTON)方法仅擅长单层替换,难以处理叠穿与去层。本文提出层级虚拟试穿(LVTON)框架,包含新基准与方法,可在保留原有穿搭的基础上实现顺序叠穿。我们发现现有VTON范式因依赖无布料特性的表征及单物品数据集,丢失了关键的叠穿上下文。核心洞察是将挑战拆解为两类能力:(1) 通用VTON先验(如形变、身份保持),(2) 特定叠穿知识(如叠穿顺序、遮挡推理)。模型首先通过自动数据生成管道(基于分割与修复的时尚视频合成)学习通用先验;随后在小型专用LVTON数据集上微调,掌握叠穿逻辑。该方法在自建的LVTON基准上达当前最优,在传统VTON基准上也表现出卓越泛化性,微调后取得新纪录,并具备零样本能力。
原文摘要 · Abstract (English)
In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer swap. This fundamental real-world task remains a challenge in existing Virtual Try-On (VTON) methods, which excel at single-layer replacement but are not designed to layer or de-layer an existing outfit. This paper proposes Layering Virtual Try-On (LVTON), a layering benchmark and method that preserves an existing outfit while enabling sequential layering. We find that current VTON paradigms are fundamentally ill-equipped for LVTON, as their reliance on cloth-agnostic representations and single-item datasets discards essential layering context. Our key insight is that the LVTON challenge must be disentangled into two distinct competencies: (1) General VTON Priors (e.g., deformation, identity preservation) and (2) Specific Layering Knowledge (e.g., layering order and occlusion reasoning). First, our model obtains general VTON priors by being trained on data produced by an automatic data generation pipeline that synthesizes samples from fashion videos via segmentation and inpainting. Second, the model is fine-tuned on a small, dedicated LVTON dataset to learn the layering logic. Our method achieves state-of-the-art results on our LVTON benchmark and demonstrates superior generalizability on traditional VTON benchmarks, setting new state-of-the-art results when fine-tuned and exhibiting zero-shot capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。