提出分阶段穿衣变形网络,让虚拟试穿更真实、肢体细节更自然。
Limb-Aware Virtual Try-On Network with Progressive Clothing Warping
- 分两阶段对齐服装与人体,逐步实现精细变形。
- 引入重力感知损失,更好处理服装边缘贴合度。
- 专注肢体区域纹理融合,提升试穿后肢体细节真实感。
基于图像的虚拟试穿旨在将店内服装图像转移到人物图像上。现有方法多采用全局单一形变直接进行服装变形,缺乏对店内服装的细粒度建模,导致服装外观失真。同时,由于使用不依赖服装的人体表征,无法有效生成肢体细节。为此,本文提出一种名为PL-VTON的肢体感知虚拟试穿网络,通过分阶段的精细服装变形生成高质量试穿结果,并保留真实肢体细节。具体而言,提出渐进式服装变形(PCW),显式建模店内服装的位置与尺寸,采用两阶段对齐策略逐步对齐服装与人体;设计重力感知损失,考虑人体穿戴服装的贴合度,优化服装边缘处理;提出人体解析估计器(PPE),使用非肢体目标解析图将人体语义分割为多个区域,提供结构约束以缓解服装与身体区域间的纹理溢出;最后引入肢体感知纹理融合(LTF),先生成粗略试穿结果,再在肢体感知引导下融合肢体纹理以细化肢体细节。大量实验表明,PL-VTON在定性和定量上均优于现有最先进方法。
原文摘要 · Abstract (English)
Image-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directly, which lacks fine-grained modeling of in-shop clothing and leads to distorted clothing appearance. In addition, existing methods usually fail to generate limb details well because they are limited by the used clothing-agnostic person representation without referring to the limb textures of the person image. To address these problems, we propose Limb-aware Virtual Try-on Network named PL-VTON, which performs fine-grained clothing warping progressively and generates high-quality try-on results with realistic limb details. Specifically, we present Progressive Clothing Warping (PCW) that explicitly models the location and size of in-shop clothing and utilizes a two-stage alignment strategy to progressively align the in-shop clothing with the human body. Moreover, a novel gravity-aware loss that considers the fit of the person wearing clothing is adopted to better handle the clothing edges. Then, we design Person Parsing Estimator (PPE) with a non-limb target parsing map to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and body regions. Finally, we introduce Limb-aware Texture Fusion (LTF) that focuses on generating realistic details in limb regions, where a coarse try-on result is first generated by fusing the warped clothing image with the person image, then limb textures are further fused with the coarse result under limb-aware guidance to refine limb details. Extensive experiments demonstrate that our PL-VTON outperforms the state-of-the-art methods both qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。