无需遮罩的虚拟试衣,用扩散模型生成高精度穿着图
MFP-VTON: Enhancing Mask-Free Person-to-Person Virtual Try-On via Diffusion Transformer
- 基于预训练扩散变换器,无须服装掩码直接生成试穿图像
- 引入焦点注意力损失,提升参考服装与目标人物细节的还原度
- 适用于真实场景下的快速虚拟试衣,适合电商与个性化推荐
服装到人物的虚拟试衣(VTON)任务旨在生成人物穿着参考服装的拟合图像,已取得显著进展。然而,获取标准服装数据往往比使用人物已穿服装更困难。为提升易用性,我们提出MFP-VTON,一种无遮罩的人物到人物虚拟试衣框架。鉴于人物到人物数据稀缺,我们改造了原有的服装到人物模型与数据集,构建出专用于该任务的数据集。方法基于预训练的扩散变换器,利用其强大的生成能力。在无遮罩模型微调过程中,引入焦点注意力损失,强调参考人物的服装区域以及目标人物服装外的细节。实验表明,该模型在人物到人物及服装到人物的VTON任务中均表现优异,生成高质量拟合图像。
原文摘要 · Abstract (English)
The garment-to-person virtual try-on (VTON) task, which aims to generate fitting images of a person wearing a reference garment, has made significant strides. However, obtaining a standard garment is often more challenging than using the garment already worn by the person. To improve ease of use, we propose MFP-VTON, a Mask-Free framework for Person-to-Person VTON. Recognizing the scarcity of person-to-person data, we adapt a garment-to-person model and dataset to construct a specialized dataset for this task. Our approach builds upon a pretrained diffusion transformer, leveraging its strong generative capabilities. During mask-free model fine-tuning, we introduce a Focus Attention loss to emphasize the garment of the reference person and the details outside the garment of the target person. Experimental results demonstrate that our model excels in both person-to-person and garment-to-person VTON tasks, generating high-fidelity fitting images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。