PROMO用高效模型实现高保真虚拟试穿,速度与质量兼得。
PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On
- 基于流匹配的DiT架构,结合多模态条件拼接和自参考机制。
- 在标准数据集上视觉保真度超越现有方法,推理速度快于同类模型。
- 适合需要快速生成高真实感试穿图的电商与设计场景。
虚拟试穿(VTON)已成为在线零售的核心功能,逼真的试穿结果能提供可靠的合身指导,减少退货,惠及消费者与商家。基于扩散模型的VTON方法虽可实现照片级合成,但常依赖复杂结构(如辅助参考网络),且采样缓慢,难以兼顾保真度与效率。本文将VTON视为需满足主体保留、纹理忠实迁移与无缝融合三个要求的结构化图像编辑问题。所提训练框架通用性强,可迁移到更广泛的图像编辑任务;此外,VTON生成的成对数据为训练通用编辑器提供了丰富监督资源。我们提出PROMO,一个基于流匹配DiT骨干网络、结合潜空间多模态条件拼接的可提示虚拟试穿框架。通过条件高效性与自参考机制,显著降低推理开销。在标准基准上,PROMO在视觉保真度上超越先前VTON方法及通用图像编辑模型,同时保持质量与速度的竞争力。结果表明,流匹配变压器结合潜空间多模态条件与自参考加速,是高质量虚拟试穿的有效且训练高效的解决方案。
原文摘要 · Abstract (English)
Virtual Try-on (VTON) has become a core capability for online retail, where realistic try-on results provide reliable fit guidance, reduce returns, and benefit both consumers and merchants. Diffusion-based VTON methods achieve photorealistic synthesis, yet often rely on intricate architectures such as auxiliary reference networks and suffer from slow sampling, making the trade-off between fidelity and efficiency a persistent challenge. We approach VTON as a structured image editing problem that demands strong conditional generation under three key requirements: subject preservation, faithful texture transfer, and seamless harmonization. Under this perspective, our training framework is generic and transfers to broader image editing tasks. Moreover, the paired data produced by VTON constitutes a rich supervisory resource for training general-purpose editors. We present PROMO, a promptable virtual try-on framework built upon a Flow Matching DiT backbone with latent multi-modal conditional concatenation. By leveraging conditioning efficiency and self-reference mechanisms, our approach substantially reduces inference overhead. On standard benchmarks, PROMO surpasses both prior VTON methods and general image editing models in visual fidelity while delivering a competitive balance between quality and speed. These results demonstrate that flow-matching transformers, coupled with latent multi-modal conditioning and self-reference acceleration, offer an effective and training-efficient solution for high-quality virtual try-on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。