无需模特图和掩码,一模型实现任意姿态的虚拟试穿与脱衣
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- 基于双向扩散机制,仅需一张人像和一件衣服输入
- 支持跨人物换装与任意姿态生成,效果优于现有方法
- 适合电商试衣、虚拟穿搭等真实场景应用
近期基于扩散模型的虚拟试衣方法在图像驱动试衣方面取得显著进展,实现了更逼真、端到端的服装合成。然而,多数现有方法仍依赖展示衣物和分割掩码,且难以处理灵活的姿态变化,限制了其在真实场景中的实用性——例如用户无法将他人穿着的衣物转移到自己身上,生成结果也通常局限于参考图像的姿态。本文提出OMFA(One Model For All),一种统一的扩散框架,可同时实现虚拟试穿与脱衣,无需展示衣物,支持任意姿态。该框架受离散扩散语言模型的掩码范式启发,构建于双向Tweedie扩散过程之上,实现潜在空间中的目标选择性去噪。不同于施加下身约束的方法,OMFA为完全无掩码设计,仅需单张人像与目标服装作为输入,支持灵活的穿搭组合与跨人物服装迁移,更契合实际使用需求。此外,通过引入基于SMPL-X的姿态条件控制,仅凭一张图像即可实现多视角、任意姿态的试衣。大量实验表明,OMFA在试穿与脱衣任务上均达到当前最优性能,提供了一种实用且通用的虚拟服装合成方案。
原文摘要 · Abstract (English)
Recent diffusion-based approaches have made significant advances in image-based virtual try-on, enabling more realistic and end-to-end garment synthesis. However, most existing methods remain constrained by their reliance on exhibition garments and segmentation masks, as well as their limited ability to handle flexible pose variations. These limitations reduce their practicality in real-world scenarios; for instance, users cannot easily transfer garments worn by one person onto another, and the generated try-on results are typically restricted to the same pose as the reference image. In this paper, we introduce OMFA (One Model For All), a unified diffusion framework for both virtual try-on and try-off that operates without the need for exhibition garments and supports arbitrary poses. OMFA is inspired by the mask-based paradigm of discrete diffusion language models and unifies try-on and try-off within a bidirectional framework. It is built upon a Bidirectional Tweedie Diffusion process for target-selective denoising in latent space. Instead of imposing lower body constraints, OMFA is an entirely mask-free framework that requires only a single portrait and a target garment as inputs, and is designed to support flexible outfit combinations and cross-person garment transfer, making it better aligned with practical usage scenarios. Additionally, by leveraging SMPL-X-based pose conditioning, OMFA supports multi-view and arbitrary-pose try-on from just one image. Extensive experiments demonstrate that OMFA achieves state-of-the-art results on both try-on and try-off tasks, providing a practical and generalizable solution for virtual garment synthesis. Project page: https://onemodelforall.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。