无需用户画图,仅用一张人像和一件衣服就能实现高保真虚拟试穿。
MF-VITON: High-Fidelity Mask-Free Virtual Try-On with Minimal Input
- 两阶段流程:先用有遮罩模型生成海量数据,再微调出无遮罩版本。
- 在多个评测集上超越现有方法,衣物纹理与轮廓保持度显著提升。
- 适合追求真实试穿效果的电商、服装设计等场景使用。
虚拟试穿(VITON)近年借助强大的文本到图像扩散模型,在图像真实感和服装细节保留方面取得显著进展。然而,现有方法普遍依赖用户提供的掩码,导致操作复杂且因输入不准确影响性能,如图1(a)所示。为此,我们提出无遮罩虚拟试穿(MF-VITON)框架,仅需单张人物图像和目标服装即可实现高保真试穿,彻底消除对辅助掩码的需求。该方法采用新颖的两阶段流程:(1) 利用已有基于遮罩的VITON模型合成高质量数据集,包含多样化的真人与服装配对样本,并通过不同背景增强以模拟真实场景;(2) 在生成数据集上对预训练遮罩模型进行微调,使新模型在无需遮罩的情况下仍能完成精准的服装迁移,同时保持服装纹理与形状的真实感。实验表明,该框架在服装转移精度与视觉真实性方面达到当前最优水平,所提无遮罩模型显著优于以往基于遮罩的方法,树立了新的基准。更多信息请访问项目页面:https://zhenchenwan.github.io/MF-VITON/。
原文摘要 · Abstract (English)
Recent advancements in Virtual Try-On (VITON) have significantly improved image realism and garment detail preservation, driven by powerful text-to-image (T2I) diffusion models. However, existing methods often rely on user-provided masks, introducing complexity and performance degradation due to imperfect inputs, as shown in Fig.1(a). To address this, we propose a Mask-Free VITON (MF-VITON) framework that achieves realistic VITON using only a single person image and a target garment, eliminating the requirement for auxiliary masks. Our approach introduces a novel two-stage pipeline: (1) We leverage existing Mask-based VITON models to synthesize a high-quality dataset. This dataset contains diverse, realistic pairs of person images and corresponding garments, augmented with varied backgrounds to mimic real-world scenarios. (2) The pre-trained Mask-based model is fine-tuned on the generated dataset, enabling garment transfer without mask dependencies. This stage simplifies the input requirements while preserving garment texture and shape fidelity. Our framework achieves state-of-the-art (SOTA) performance regarding garment transfer accuracy and visual realism. Notably, the proposed Mask-Free model significantly outperforms existing Mask-based approaches, setting a new benchmark and demonstrating a substantial lead over previous approaches. For more details, visit our project page: https://zhenchenwan.github.io/MF-VITON/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。