arXiv:2504.13078cs.CVcs.AI2025-04ICCV被引 8

MGT让虚拟试衣扩展到多件衣物场景,自动从真人照片中提取服装特征。

MGT: Extending Virtual Try-Off to Multi-Garment Scenarios

  • 基于扩散模型与SigLIP图像条件,从真人照片中提取多类服装特征。
  • 在VITON-HD上达到当前最佳性能,对上衣、下装和连衣裙均有效。
  • 可配合试衣模型减少肤色等无关属性传递,适合服装设计与电商应用。

计算机视觉正在通过虚拟试衣(VTON)和虚拟试穿(VTOFF)重塑时尚产业。VTON利用目标服装图和人物照片生成穿着效果,而更具挑战性的跨人虚拟试衣(p2p-VTON)则使用他人穿着的服装图。相比之下,VTOFF能从穿着者照片中提取标准化服装图像。我们提出多服装试穿扩散模型MGT,基于潜在扩散架构,结合SigLIP图像条件以捕捉服装的形状、纹理和图案特征。为应对服装多样性,MGT引入类别特定嵌入,在VITON-HD上取得当前最优表现,并在DressCode上表现竞争力。当与VTON模型结合时,可有效减少肤色等非人物特征的传递,确保保留个体特异性。演示、代码与模型已公开:https://rizavelioglu.github.io/tryoffdiff/

原文摘要 · Abstract (English)

Computer vision is transforming fashion industry through Virtual Try-On (VTON) and Virtual Try-Off (VTOFF). VTON generates images of a person in a specified garment using a target photo and a standardized garment image, while a more challenging variant, Person-to-Person Virtual Try-On (p2p-VTON), uses a photo of another person wearing the garment. VTOFF, in contrast, extracts standardized garment images from photos of clothed individuals. We introduce Multi-Garment TryOffDiff (MGT), a diffusion-based VTOFF model capable of handling diverse garment types, including upper-body, lower-body, and dresses. MGT builds on a latent diffusion architecture with SigLIP-based image conditioning to capture garment characteristics such as shape, texture, and pattern. To address garment diversity, MGT incorporates class-specific embeddings, achieving state-of-the-art VTOFF results on VITON-HD and competitive performance on DressCode. When paired with VTON models, it further enhances p2p-VTON by reducing unwanted attribute transfer, such as skin tone, ensuring preservation of person-specific characteristics. Demo, code, and models are available at: https://rizavelioglu.github.io/tryoffdiff/

虚拟试衣扩散模型服装生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。