用户画线就能自由调整虚拟试衣效果,让穿搭更灵活多样。
MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On

- 用草图生成双区域掩码,精准引导服装布局
- 分上下身独立调整空间结构,实现精细适配
- 适合想自由定制穿搭风格的设计师和用户
虚拟试衣旨在将商店内拍摄的服装图像贴合到特定人体上。理想的虚拟试衣方法应提供多样且灵活的穿搭选项,准确反映真实场景中多样的穿着风格,满足个人偏好与时尚追求。然而,现有方法大多直接替换服装并沿用原有穿搭模式,控制力有限,导致试衣结果单一单调。为探索虚拟试衣中的更多时尚可能,本文提出MOFA-VTON,通过用户绘制的曲线草图实现对试衣结果的细粒度调整。首先设计一种掩码构建策略,将用户草图转化为双区域掩码,替代传统无服装感知的掩码,为生成过程提供细粒度布局指导;其次引入布局调整模块,利用交叉注意力机制分别学习人体上下半身的布局对应关系,优化两部分的空间排布。该方法实现了目标服装的灵活与精细适配,突破固定布局限制。在VITON-HD和DressCode数据集上的大量实验表明,MOFA-VTON优于现有最先进方法,显著提升虚拟试衣的时尚多样性。
原文摘要 · Abstract (English)
Virtual try-on aims to fit an in-shop clothing image onto a specific human body. An optimal virtual try-on method should provide diverse and flexible dressing options, accurately reflecting the varied wearing styles encountered in real-life scenarios, tailored to individual preferences and fashion aspirations. However, current methods predominantly perform a direct replacement of the original clothing with the target clothing, following the same dressing pattern. This limited control over clothing adaptation may result in fixed and monotonous try-on outputs. To delve into More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On, we propose a novel virtual try-on method, termed MOFA-VTON, which allows adjustment for clothing adaptations in try-on results through simple sketches by users. Specifically, we first design a mask construction strategy that transforms user-drawn curve sketches into a dual-region mask, replacing the traditional clothing-agnostic mask and providing fine-grained layout guidance for the subsequent generation process. Further, we propose layout adjustment blocks that utilize the cross-attention mechanism to independently learn layout correspondences for upper and lower regions of the human body, refining the spatial arrangement of the two regions. With these implementations, our method enables flexible and fine-grained adaptations of target clothing, overcoming the constraints of a fixed layout. Extensive experiments on VITON-HD and DressCode datasets demonstrate that our proposed MOFA-VTON outperforms previous state-of-the-art methods and provides more fashion possibilities for virtual try-on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。