无需训练即可跨场景通用试穿,支持多人体真实换装。
OmniVTON: Training-Free Universal Virtual Try-On
- 分离服装与姿态条件,解耦生成过程
- 实现跨域高保真换装,支持多人体场景
- 首个无需训练的通用试穿框架,适合实际应用
基于图像的虚拟试穿技术通常依赖监督式店内方法(保真度高但泛化差)或无监督户外方法(适应性强但受数据偏差限制)。现有方法难以兼顾跨场景通用性。本文提出OmniVTON,首个无需训练的通用虚拟试穿框架,通过解耦服装与姿态条件,在多样场景下实现纹理保真与姿态一致。为保留服装细节,引入服装先验生成机制并采用连续边界拼接技术;为精准对齐姿态,利用DDIM反演提取结构信息并抑制纹理干扰。该设计有效消除扩散模型在多重条件下的固有偏见。实验表明,OmniVTON在多个数据集、服装类型和应用场景中表现优异,首次实现单场景多人体真实换装。代码已开源。
原文摘要 · Abstract (English)
Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve fine-grained texture retention. For precise pose alignment, we utilize DDIM inversion to capture structural cues while suppressing texture interference, ensuring accurate body alignment independent of the original image textures. By disentangling garment and pose constraints, OmniVTON eliminates the bias inherent in diffusion models when handling multiple conditions simultaneously. Experimental results demonstrate that OmniVTON achieves superior performance across diverse datasets, garment types, and application scenarios. Notably, it is the first framework capable of multi-human VTON, enabling realistic garment transfer across multiple individuals in a single scene. Code is available at https://github.com/Jerome-Young/OmniVTON
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。