arXiv:2508.17614cs.CV2025-08被引 5

无需人体遮罩的虚拟试穿,通过多模态扩散模型实现精准控制

JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

  • 用多模态扩散变压器融合人像与服装图像,直接控制试穿效果
  • 在DressCode数据集上显著超越现有方法,真人评价也更优
  • 适合需要真实场景试穿和精细属性调节的研究与应用

虚拟试穿系统长期受限于对人工人体遮罩的依赖、服装属性控制精度不足,以及在真实场景下的泛化能力差。本文提出JCo-MVTON(面向无遮罩虚拟试穿的联合可控多模态扩散变压器),通过将基于扩散的图像生成与多模态条件融合结合,克服上述问题。该框架基于多模态扩散变压器(MM-DiT)骨干网络,通过专门的条件路径,在去噪过程中直接融合参考人像与目标服装图像特征,并在自注意力层内进行特征融合,结合优化的位置编码与注意力掩码,实现精确的空间对齐与服装-人体整合。为解决数据稀缺与质量低的问题,提出双向生成策略构建数据集:一管道使用基于遮罩的模型生成逼真参考图像,另一对称的“脱衣”模型以自监督方式恢复对应服装图像。合成数据经严格人工筛选,支持视觉保真度与多样性的迭代提升。实验表明,JCo-MVTON在DressCode等公开基准上达到当前最优性能,量化指标与人类评估均优于现有方法,且在真实场景中表现强于商用系统。

原文摘要 · Abstract (English)

Virtual try-on systems have long been hindered by heavy reliance on human body masks, limited fine-grained control over garment attributes, and poor generalization to real-world, in-the-wild scenarios. In this paper, we propose JCo-MVTON (Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-On), a novel framework that overcomes these limitations by integrating diffusion-based image generation with multi-modal conditional fusion. Built upon a Multi-Modal Diffusion Transformer (MM-DiT) backbone, our approach directly incorporates diverse control signals -- such as the reference person image and the target garment image -- into the denoising process through dedicated conditional pathways that fuse features within the self-attention layers. This fusion is further enhanced with refined positional encodings and attention masks, enabling precise spatial alignment and improved garment-person integration. To address data scarcity and quality, we introduce a bidirectional generation strategy for dataset construction: one pipeline uses a mask-based model to generate realistic reference images, while a symmetric ``Try-Off'' model, trained in a self-supervised manner, recovers the corresponding garment images. The synthesized dataset undergoes rigorous manual curation, allowing iterative improvement in visual fidelity and diversity. Experiments demonstrate that JCo-MVTON achieves state-of-the-art performance on public benchmarks including DressCode, significantly outperforming existing methods in both quantitative metrics and human evaluations. Moreover, it shows strong generalization in real-world applications, surpassing commercial systems.

虚拟试穿扩散模型多模态无遮罩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。