arXiv:2606.29319cs.CV2026-06中稿 · ECCV

仅用6步生成高保真无遮挡试穿图像,突破传统方法依赖掩码和多步采样的限制。

FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On

论文配图:FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On
图 1 · 摘自论文原文
  • 引入噪声优化与指令驱动模块,实现无掩码条件下的高效试穿生成。
  • 6步采样即达到30步基线的图像质量,显著提升推理效率。
  • 适用于追求快速、真实试穿体验的电商与虚拟服装应用。

基于图像的虚拟试穿(IVTON)借助扩散模型取得显著进展,但现有方法需大量采样步骤且依赖代价高昂的掩码与辅助网络。此外,缺乏大规模无掩码配对数据集也制约了无掩码IVTON的发展。本文提出FDM-MFVT,一种面向无掩码虚拟试穿的少步扩散模型,集成外观感知噪声优化模块(OANO)与指令驱动试穿模块(IDT),以提升效率与灵活性。OANO模块利用输入图像初始化噪声空间,仅需6步即可生成高保真试穿图像,优于30步基线。IDT模块通过虚拟试穿提示与高效适配,仅凭衣物与人物图像生成高质量结果。我们还构建了包含30,000对样本的无掩码IVTON数据集MFVT。实验表明,FDM-MFVT在少步推理下,量化与定性指标均优于基于掩码与无掩码的基线方法。

原文摘要 · Abstract (English)

Image-based Virtual Try-On (IVTON) has greatly advanced through diffusion models, yet existing methods require many sampling steps and depend on masks with costly auxiliary networks. In addition, the absence of large-scale mask-free paired datasets further limits the development of mask-free IVTON. We propose FDM-MFVT, a few-step diffusion model for mask-free IVTON, integrating an Outfit-aware Noise Optimization Module (OANO) and an Instruction-driven Try-on Module (IDT) to enhance efficiency and flexibility.The OANO module initializes the alignment space with noise using the input image and only needs 6 steps to generate a higher-fidelity try-on image compared to 30 steps.The IDT module uses virtual try-on prompts and efficient adaptation to generate high-quality results from garment and person images alone. We further introduce MFVT, a 30,000-pair mask-free IVTON dataset. Experiments show that FDM-MFVT achieves superior quantitative and qualitative results with fewer inference steps than mask-based and mask-free baseline methods.

虚拟试穿扩散模型少步采样无掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。